Search
Every feed is a document in OpenSearch, at OPENSEARCH_URL. xixo indexes a feed
again whenever the feed or one of its references is saved, and whenever an
analysis of it finishes. It removes the document when the feed is destroyed.
The index is a derived copy. Postgres holds the records, and a search returns ids that xixo loads
from Postgres in ranked order.
What a document holds
Section titled “What a document holds”| Field | From | Analyzed as |
|---|---|---|
tenant_id |
the feed’s tenant | long |
type, mime |
the feed | keyword |
tags |
the keys of its tags, across the feed and its children | path, and tags.raw as keyword |
key, title |
the feed | path |
locator_key |
its original references, joined | path |
note |
the feed | standard |
summary |
every summary across the feed and its children | standard |
body |
everything its analyses extracted, across the feed and its children | standard |
resource_ids |
the resources where its original references live | long |
created_at |
the feed | date |
embedding |
the feed’s vector, when it has one | knn_vector, HNSW, cosine |
The path analyzer splits on slashes, backslashes, dashes, underscores, dots, and whitespace, and
then lowercases, so a search for acme finds invoices/2026/acme-q3.pdf. Only original references
contribute a locator or a resource id. A preview or a thumbnail never does. The body leaves out the
summary step, which has its own field, and the bookkeeping steps, such as placement and derived,
whose results are paths and store names. The mapping is not dynamic, so a field that is not listed
here is not indexed.
Lexical search
Section titled “Lexical search”A query with text runs a multi_match over these fields, with these boosts:
| 3 | title, key, and tags |
| 2 | note and summary |
| 1 | locator_key and body |
title, note, summary, and body are also indexed a second time through the kstem stemmer,
and searched with the same boosts, so a word finds its other forms: invoices finds invoice, and
kilners finds Kilner.
The first search is strict. It uses the and operator, so every word has to match, and the
bool_prefix type, so the last word matches as a prefix. A query of acme inv finds a feed that
contains acme and invoice.
If the strict search does not fill the requested page, and the query has three or more words, xixo runs a second, looser search. The loose search requires 60% of the words to match, or as few as the caller asks for: an answer’s evidence asks for any one of the question’s keywords. Its results go after the strict results, without duplicates, and the reported total is the larger of the two totals. A query of one or two words gets only the strict search.
A query with no text matches everything. Facets narrow the results in either case, and each facet takes one value or a list:
| Facet | Field |
|---|---|
type |
type |
mime |
mime |
tag |
tags.raw |
Semantic search
Section titled “Semantic search”Semantic search is added when all of these are true:
- the query has text;
- an inference resource declares a model for the
embeddingrole; - the requested page ends within the first 200 results.
xixo embeds the query and caches the vector for a day per tenant, resource, model, and text. It then
asks for the 200 nearest neighbors under the same facets. The neighbors are cut to the ones whose
cosine similarity is at least XIXO_SEMANTIC_FLOOR (0.55 by default), and within 0.12 of the best
match. A query whose meaning matches nothing well therefore adds no semantic results.
When no facet is set, xixo also asks the passage index for the 200 passages nearest the query, cut
the same way, and ranks their feeds by their best passage. This finds a long document by something
said deep inside it, which its own vector, drawn mostly from its opening, does not carry. The
search tool gives each result found this way the passage and where it starts.
The top 200 lexical results and the remaining neighbors, of feeds and of passages, are fused by
reciprocal rank fusion. Each list adds 1 / (60 + rank) to a feed’s score, with ranks counted from one, and the fused list is
sorted by score. Fusion uses only the rank in each list. The two engines score on different scales,
so their scores are never compared.
Search is lexical only when the page is past the first 200 results, when no resource serves the
embedding role, or when embedding the query fails. If the engine refuses the vector query, xixo
logs the error and the vector query adds nothing. Pages use offsets: from and limit select a
slice of the fused list, and the reported total is the larger of the lexical total and the length
of the fused list.
Listing without the index
Section titled “Listing without the index”The feeds query does not use OpenSearch. It lists feeds from Postgres, newest first, and pages by
id with after. Its arguments narrow the list:
type, types |
one type, or any of several |
mime |
feeds with a reference of that mime type |
resourceId |
feeds with a reference in that resource |
tag |
feeds connected to that tag |
connectedTo |
feeds connected to the feed with that id, in either direction |
topLevel |
only feeds that were not extracted from another feed |
When a feed is selected in the catalog, the catalog uses connectedTo to list the feeds connected to it.
Embeddings
Section titled “Embeddings”EmbedItemsJob runs every minute in every tenant and embeds up to fifty feeds that have no
embedded_at. The text it embeds is the feed’s title, tags, summaries, note, and the first 2,000
characters of its body, up to 8,000 characters in total. xixo stores a digest of that text and the
model name next to the vector. If the digest has not changed, xixo sets embedded_at without
embedding the feed again. Changing a title or a note, finishing an analysis, or saving a reference
clears embedded_at.
Some embedding models were trained with a prefix on each text, one for what is stored and one for
what is searched, and match poorly without it. xixo puts search_document: and search_query:
before texts for nomic-embed-text, passage: and query: for e5 models, and the searching
instruction mxbai, bge, and snowflake-arctic-embed expect before queries. A model backend’s Query
prefix and Document prefix replace these.
The model and its prefixes make a signature. Every digest includes it, and the backend keeps the
signature its vectors were made with. When the two differ, because the model or a prefix changed,
the next sweep clears embedded_at on every feed and passage in the tenant, since vectors made two
ways cannot be compared. The vector’s dimension is XIXO_EMBEDDING_DIMENSIONS,
768 by default.
Passages
Section titled “Passages”A feed whose text is longer than 1,500 characters is also cut into passages of up to 1,500
characters. A passage never crosses into the next section of the feed’s outline: a sheet of a
workbook, a markdown heading, a page of a PDF, or a minute of a transcript. Within a section, each
passage ends at the first of a paragraph, a line, a sentence, a clause, or a word found in its last
40 percent, and the next starts 200 characters before it ends, so a clause across a boundary is
whole in one of them. A passage that ends at a section starts the next one there. The passages table holds each passage’s text, where it starts and ends, and its
vector. A digest of the text in feeds.passages_digest means a feed is cut again only when its
text changes.
The same job cuts feeds that have no passages digest, up to fifty a minute, and embeds up to 64
passages a minute, each with its feed’s title and the name of its section before it, such as
elm-street-pantry-plan.xlsx › Buy List. The vectors are indexed in their own
index, _passages after the feeds’ index name, and every query on it filters by tenant_id.
The feed tool’s find returns a long document’s passages that contain the words asked for, ranked
by how many of them each passage holds, together with the passages nearest in meaning. Each says
whether it matched by words, by meaning, or both.
Tenants
Section titled “Tenants”All tenants’ documents share one index. Each tenant searches through its own alias, created with the
tenant, which filters on tenant_id. The nearest-neighbor query repeats that filter inside itself.
xixo refuses a search that has no tenant.
Rebuilding
Section titled “Rebuilding”xixo hashes the index’s settings and mapping into a stamp and stores it in the index’s metadata.
RebuildSearchIndexJob runs daily at 3 a.m. and does nothing unless the stamp has changed or the index
is missing. When either happens, the job builds a new versioned index and reindexes every tenant into
it in pages of 200, under a run of kind reindex. It refuses to promote the new index if it holds
fewer documents than there are feeds. Promotion moves the main alias and every tenant alias in one
request, and deletes the old index. If promotion fails, the job deletes the new index and the old
one keeps serving. Feeds that changed during the rebuild are indexed again afterward.