Search your app's data
Task: find entries of a collection by the words in them — messages that mention “invoice”, documents whose title starts with “roa” — without reading every entry in WASM.
A scan in WASM costs gas for every entry it reads. On search-chat, a
lowercase substring test over every message costs about 63,000 gas a message,
so it runs out of the default one-billion budget before 16,000 messages. An
indexed query costs the same at 2,000 or 200,000 messages: what it reads is
the page of hits, never the collection.
Each node keeps a full-text index (tantivy) of each context’s indexed collections, in its own store, built from the state it holds. The index never enters a delta, a snapshot or the root hash.
Opt in
Section titled “Opt in”-
Derive
Searchableon the collection’s value type and mark the fields to index:use calimero_sdk::app;use calimero_sdk::borsh::{BorshDeserialize, BorshSerialize};use calimero_storage::collections::{LwwRegister, UnorderedMap};#[derive(BorshSerialize, BorshDeserialize, app::Mergeable, app::Searchable)]#[borsh(crate = "calimero_sdk::borsh")]pub struct Message {#[search(keyword)]pub sender: LwwRegister<String>,#[search(text, infix)]pub text: LwwRegister<String>,#[search(number)]pub ts: LwwRegister<u64>,}Attribute Indexed as Queried by #[search(text)]Words: Unicode segmentation, case and accent folding, CJK bigrams. BM25-ranked. weight = Nboosts the field (hundredths, default 100)words,prefix,fuzzy#[search(text, infix)]Words, and trigrams also substring(3 characters or more)#[search(keyword)]One exact value, not tokenized or scored Query::eq#[search(number)]A u64Query::rangename = "..."Renames the field in the index with = pathWhat path(&field)returns (anOption<String>for text or keyword, anOption<u64>for a number) instead of the field itselfas its kind A text or keyword field’s type implements
SearchText(String,&str,Option<T>,LwwRegister<T>, and the collaborative textsFugueText,RichTextandRichDocument, which index the text a reader sees, formatting left out); a number field’sSearchNumber(u8tou64,Option<T>,LwwRegister<T>). A field without#[search]is not indexed.withis for a value whose stored form is not what a reader searches: editor HTML, say, indexed as its plain text.fn plain_text(html: &LwwRegister<String>) -> Option<String> { /* drop the tags */ }#[derive(BorshSerialize, BorshDeserialize, app::Mergeable, app::Searchable)]#[borsh(crate = "calimero_sdk::borsh")]#[search(index_if = Reply::shown)]pub struct Reply {#[search(text, with = plain_text)]pub html: LwwRegister<String>,#[search(number)]pub ts: LwwRegister<u64>,pub hidden: LwwRegister<bool>,}impl Reply {fn shown(&self) -> bool { !*self.hidden.get() }}#[search(index_if = path)]on the struct keeps a value out of the index whilepath(&value)is false — a soft-deleted or hidden entry. It leaves the index on the write that hides it and comes back on the one that shows it, andcollection.searchnever returns it in between. -
Name the collections of your state that are indexes:
#[app::state]pub struct Chat {messages: UnorderedMap<String, Message>,replies: AuthoredVector<Reply>,pinned: Moderated<SortedMap<String, Reply>>,}app::search_indexes!(Chat {"messages" (version = 1) => messages,"replies" (version = 1) => replies | pinned,});The macro generates the three exports the node’s indexer calls (
__calimero_search_schema,__calimero_search_extract,__calimero_search_scan). Any collection whose value isSearchablecan be an index:Collection A hit’s keyUnorderedMap<K, V>,SortedMap<K, V>,IndexedMap<K, V>The entry’s key Guarded<C, P>over one of them (Authored,Moderated,WriteOnce, …)The entry’s key; an owned entry is found under its owner AuthoredVector<V>The entry’s id; position_of_idgives its position when you need onea | bmakes one index over several collections of the same value type: one ranking, one cursor, a hit from whichever holds it. Find which with each collection’ssearch_entry(hit.id), as below.The version is yours to bump when a field’s meaning changes without its schema changing (a new
withfunction, say); the node rebuilds the index. A changed schema rebuilds it on its own. -
Query from a view:
use calimero_sdk::search::{Query, SearchCollection};#[app::view]pub fn search(&self, text: String, sender: Option<String>) -> app::Result<Vec<String>> {let mut query = Query::words(text).limit(20);if let Some(sender) = sender {query = query.eq("sender", sender);}let results = self.messages.search("messages", &query)?;Ok(results.hits.into_iter().map(|hit| hit.key).collect())}Over an index of several collections, query with
Query::runand read each hit back from the collection that holds it:let response = Query::words(text).newest_first("ts").limit(20).run("replies")?;for hit in &response.hits {let reply = match self.replies.search_entry(hit.id)? {Some((_, reply)) => reply,None => match self.pinned.search_entry(hit.id)? {Some((_, reply)) => reply,None => continue, // gone since the index saw it},};// ...}
That is the whole opt-in. An app that does not use search_indexes! exports
none of the three functions; for it the node writes no dirty row, opens no
index and runs nothing. It pays one export lookup per execution.
apps/search-chat is the complete example.
The query API
Section titled “The query API”| Builder | Matches |
|---|---|
Query::words(text) |
Every word of text, ranked by BM25 over the text fields |
Query::prefix(text) |
The same, with the last word as a prefix (search as you type) |
Query::substring(text) |
text inside an infix field, case- and accent-insensitive; 3 characters or more |
Query::fuzzy(text) |
Every word within one edit |
Query::all() |
Every document; pair it with filters |
.eq(field, value) |
Only hits whose keyword field is value |
.range(field, min, max) |
Only hits whose number field is in min..=max |
.newest_first(field), .oldest_first(field) |
Order the hits by a number field instead of by relevance: the newest messages that match |
.cursor(n), .limit(n) |
Paging: start after n hits; at most n hits (20 by default, 100 at most) |
collection.search(index, &query) returns Results { hits, total, next_cursor, stale }.
Each Hit { key, value, score, snippet } carries the entry as it is in state
now: the index returns entity ids, and the collection reads each one back.
A hit whose entry is gone, or that index_if now keeps out, is dropped and
counted in stale. snippet is a <b>-highlighted fragment of the text
field the words were found in (the body, when a document’s title does not
match). Under newest_first / oldest_first every score is 0. Query::run
returns the raw response (ids, no read-back) when you need it.
How it works
Section titled “How it works”Search adds one crate, calimero-search, and two node-local columns to the
store. State and the index are kept apart: state is replicated and hashed into
the root, the index is derived from it on each node and never leaves it.
┌────────────────────────── one node ──────────────────────────┐ frontend ───▶│ JSON-RPC ──▶ context manager ──▶ WASM app │ │ │ writes / views │ peers ──────▶│ sync ─────────────┘ __calimero_search_* │ │ (deltas, snapshots, ▲ │ │ │ repair; never the index) │ │search_query│ │ │ one batch │extract/ ▼ │ │ ▼ │scan search service│ │ RocksDB: State │ SearchDirty │ SearchIndex ◀──┘ │ │ │ │ ▲ └────────────────┘ │ │ ▼ │ │ │ indexer ────────┘ (background) │ └──────────────────────────────────────────────────────────────┘| Part | Where | Role |
|---|---|---|
| Generated exports | your app (search_indexes!) |
__calimero_search_schema (fields), __calimero_search_extract (entries by id), __calimero_search_scan (every entry, a page at a time). Read-only. |
| Dirty log | SearchDirty column |
One row per committed execution of a search-enabled app: state root before and after, and the ids it touched with every entry above them. |
| Index | SearchIndex column |
tantivy files, per (context, index name), in 64 KiB chunks. Each commit records the dirty-log position and state root it reflects. |
| Indexer | calimero-search, background |
Replays dirty rows through extract, or rebuilds from scan; commits every 250 ms by default. |
| Search service | calimero-search, behind search_query |
Answers a view’s query against its own context’s index. |
Opting in is detected, not configured. The node looks for
__calimero_search_extract. An app without search_indexes! has none of the
three exports, so it gets no dirty row, no index and no indexer work; only the
collections named in the macro are ever read.
A write
Section titled “A write”Replay continues only while each row’s before root equals the root the
index reached. When it doesn’t, or the context’s current root differs from the
index’s with no row left, state moved by another route (snapshot join,
repair, migration, a run with search off). The indexer then rebuilds the
index from scan, and the previous commit keeps answering until the new one
is in place.
A search
Section titled “A search”The read-back is what makes a node-local, slightly lagging index safe to show: a hit is only what your collection returns for that id right now.
What it guarantees
Section titled “What it guarantees”- A view searches only its own context. The host function takes no context argument: the node binds every call to the context the view runs in, and index keys are prefixed by the context id.
- Only views search. The node hands the search handle to read-only runs alone; a mutating method that calls it traps and commits nothing. A write that branched on a node-local index would record a decision other nodes could not reproduce.
- A hit is current state, filtered by your code. The index lags state by up to one commit interval (250 ms by default); the read-back means a deleted entry is never returned and an edited one shows its new value. Anything your collection would not return for an id (another collection’s entry, a forged id) is not a hit.
- The index follows state however state arrived. A local write and a peer’s delta stage a dirty row in the same write batch as the change, so a crash loses both or neither. Every row records the state root before and after, and every index commit the root it reflects. When state moves without a row — a node that joined by snapshot, a HashComparison or level-wise repair, a migration, writes made while the node ran search off — the chain breaks or the roots disagree, and the node rebuilds the index from a scan of state. It checks on every search view, and every 30 seconds for each context it indexed since start.
- An edit inside an entry re-indexes the entry. A dirty row names every
entry above a changed row, so typing into a
RichDocumentnested in a document’s record re-indexes that document. A nested row’s removal names no ancestors: it reaches the index with the entry’s next write, or the next rebuild. - Deleting a context deletes its index.
What “authorized” means is up to your read path: the index holds whatever
this node’s state holds for the context, and every member’s node holds the
same state. If your app keeps per-user visibility inside one context, enforce
it in the view that returns hits — including whether to return snippet.
What it costs
Section titled “What it costs”A query pays gas for the work the node does outside the guest, on top of the view’s own operators, charged when the call returns:
gas = 8,000 + 25 × documents matched + 18,500 × hits returned + 125 × response bytesThe constants turn measured host time into the gas the same time costs the
guest (about 1.76 gas per nanosecond on the reference machine); the fit and
the numbers behind it are in tools/search-bench/README.md. A refused query
(an unknown field, a substring under 3 characters) pays the fixed cost. One
execution may make at most 32 calls. Views never replicate, so this gas,
which depends on this node’s index, never has to agree between nodes.
Writes pay nothing extra in gas: the dirty row is host bookkeeping after the run, so a node with search on and a node with it off charge a write the same, which is required, since every node must agree whether a write ran out of gas. The row is 109 bytes plus 32 per changed entity id, in the same batch.
Measured on search-chat with the default budget (tools/search-bench/README.md
has the rest):
| 2,000 messages | 10,000 | 50,000 | 200,000 | |
|---|---|---|---|---|
one post, search off / on |
2.50 M / 2.50 M | 2.53 M / 2.53 M | 2.59 M / 2.59 M | 2.67 M / 2.67 M |
search view, common word, top 20 (host work included) |
2.03 M | 2.10 M | 2.27 M | 2.82 M |
search view, no match |
0.08 M | 0.08 M | 0.08 M | 0.08 M |
in-WASM scan (scan_search) |
160 M | 634 M | exhausts 1,000 M | exhausts 1,000 M |
At most 636 messages fit in one post_many call, search on or off: about
0.73 M gas a call plus 1.56 M an insert.
Limits
Section titled “Limits”- One index per
(context, index name), over one or more collections of one value type. - A query is at most 256 bytes and 4 KiB encoded; a page at most 100 hits; a cursor at most 10,000 deep.
- Fuzzy matching is edit distance 1 on words: no stemming, no synonyms, no phonetics. Substring search needs 3 characters.
- Scores are per context (each index has its own term statistics).
- Two nodes that have not converged answer differently: each indexes the state it holds.
- A schema change (a field added, renamed, retyped, reordered) or a
versionbump rebuilds the index from state: about 0.1 ms a document, while the previous commit keeps answering.
Operating it
Section titled “Operating it”Search is on by default. [context.search] in config.toml:
| Key | Default | |
|---|---|---|
enabled |
true |
Off: no dirty row, no indexer, search views trap |
commit_interval_ms |
250 |
The lag between a write and a search that finds it |
cache_mib |
32 |
Index chunks cached across all contexts |
max_open_indexes |
256 |
Open indexes (about 1.6 MiB each); least recently used idle ones close past it |
idle_close_secs |
600 |
An index unused this long closes |
audit_interval_secs |
30 |
How often indexed contexts are checked for state that moved without a row |
compact_after_mib |
64 |
Deleted index bytes after which a context’s slice of the index column is compacted |
The index lives in two RocksDB column families, SearchIndex and
SearchDirty, encrypted at rest with the rest of the store. A node upgraded
to this version creates them on start and builds each index the first time a
search view runs; nothing needs migrating. A binary that predates them will
not open a store that has them, like any other column added before.