Skip to content

Analysis

Each time xixo reads a feed, it writes one Analysis record. The record holds the cause, the status, each step’s result, every turn the model took, and a log. A feed’s current analysis is its last settled one. A new analysis starts with a copy of the current analysis’s steps, so xixo does not repeat work whose result is still valid.

Cause Raised by
upload an upload, a note, or a fetched URL, as described in adding things
sync a resource sync, a snapshot, an extracted child, or a parent whose children have all finished
keep the resource tool’s keep, for one object that was not yet analyzed
schedule an address feed’s schedule coming due
manual runFeed, analyzeFeed, the feed tool’s analyze, or a job started without an analysis
ask askCatalog, and analyzeFeed on a note that was asked, as described in asking

The status is queued or running while the analysis is open. It ends as done or failed, which count as settled, or as cancelled or gated, which do not. When an analysis settles, xixo marks the feed’s embedding as stale and indexes the feed again for search. Each log line and each status change is published on analysisProgressed.

cancelAnalysis moves a queued or running analysis to cancelled, and the job stops at its next check. It returns cancelled: false for an analysis that had already ended. The item page shows Stop beside an analysis in progress.

An analysis gets a deadline when it opens: XIXO_RUN_DEADLINE_HOURS from then, 6 by default, or no deadline when it is 0. This deadline applies only while the analysis is queued. ExpireOverdueJob runs every ten minutes and cancels a queued analysis past its deadline, with the error deadline passed. When the job starts the analysis, it replaces the deadline with the start time plus the feed’s timeout, as described in after the analyzer.

The log stops growing after 256,000 characters, and each line is truncated to 2,000. A turn stores its request truncated to 40,000 characters and the model’s reply truncated to 4,000.

A step is a named block. Its result is saved in steps, with the times it started and finished. If the block raises an error, the step saves the error and no result, so the next analysis runs it again. A saved result is reused and logged as cached, unless one of these is true:

  • the step was forced;
  • the step’s after time is later than its finished_at. The summary step sets after to the later of the last prompt change and the inference resource’s last update;
  • the reference the step read has a changed_at later than the step. When a file is stored again at the same path, this makes the steps that described its old bytes run again.

A feed’s text is what its source says, and nothing a model wrote about it. It is the results of five steps in this order: text, ocr, transcript, conversation, and place. Search indexes it as the body, passages are cut from it, and the feed tool reads and searches it. The summary, metadata, and every other step are left out of it.

What a model saw in a picture or a video is kept apart, as what the feed was described as. An image’s caption step holds the vision model’s description, and a video’s scenes step holds a caption for each of its sampled frames. Search finds a feed by them along with its summary, and an answer reads them under the heading “As a model described what it shows”.

Four steps are bookkeeping: placement, derived, answer, and drew_on. They record where a file went, which previews were rendered, what the agent said, and which feeds an answer used.

A feed’s details field turns the steps of its current analysis into labelled rows, grouped by the step that produced them, and the item page shows them under Details. Nested values are flattened into one row each, a list of records such as attachments or events numbers each record, and a row an earlier step already showed is not repeated. Embedded metadata comes last. Text, captions, scenes, tables, the summary, and the bookkeeping steps are left out.

Analyzer.for picks an analyzer by the feed’s mime type. It walks Analyzer.all in order and takes the first analyzer that claims the mime type:

Pdf, Image, Media, Page, Doc, Epub, Xlsx, Calendar, Pkpass, Email, Entry, Contact, Data, Text, Fallback

The order matters. text/calendar and text/csv are both text types. They reach Calendar and Data only because those come before Text. Fallback claims everything else. It records the size, detects the format from magic bytes, reads the file as text if it looks printable, and lists the entries if the file is a zip, a tar, or a tar compressed with gzip. A tar is known by its ustar header and a gzip by the tar inside it. The listing leaves out the pax headers and ._ files that a tar made on a Mac carries. A gzip that holds anything other than a tar is not listed.

Every analyzer that reads a file as text decodes it the same way. A UTF-16 file with a byte order mark is read as UTF-16, a UTF-8 byte order mark is dropped, and valid UTF-8 is kept as it is. Anything else is read as Windows-1252, the superset of Latin-1 that older text files are almost always in.

Pdf reads the text layer with pdftotext. When that yields fewer than 40 characters a page, the PDF is taken to be a scan: its first 30 pages are rendered at 200 dpi and read with tesseract, and the log says how many pages were read that way.

Xlsx writes every sheet into the text, each under its name, and records under outline where each sheet starts. Xlsx, and Data for a CSV or TSV file, also keep each sheet’s rows under the step tables, up to 5,000 rows a sheet. The header is the widest of a sheet’s first ten rows, so a title above it is left out. An answer works totals and counts out over these rows. A CSV or TSV file is summarized from its tables’ shape: the columns, two example rows, the values each text column holds most often, and the range of each number column and of a column whose values are mostly different, such as a date. Its rows never reach the summary, so a long ledger summarizes as quickly as a short one.

An EPUB is a zip too, and Epub claims it first. It follows META-INF/container.xml to the package document and records the title, creators, publisher, date, language, subjects, and chapter paths under the step book. The text step starts with those details and continues with each chapter’s text in reading order, one line per paragraph. A chapter listed in META-INF/encryption.xml is skipped, so a book sold with DRM is cataloged by its details alone.

Image asks exiftool for a photo’s GPS coordinates and records them under the step location. exiftool also reads the GPS block of a raw photo, such as a Nikon NEF. A GPS block with no position in it, which some cameras write on every frame, reads as latitude 0 and longitude 0, and xixo treats it as no location. When a Places resource has Name where photos were taken on, Image asks that resource for the address, records it under the step place, and tells the summary where the photo was taken. The vision model’s description of the image is its summary, and Image keeps it under caption too. A GIF, WebP, or APNG records how many frames it has under the step frames. With more than one, the vision model is sent frames spread evenly across it, in order, and told it is looking at an animation, so the description covers what happens in it. How many frames it is sent is set by two settings, four by default.

Media claims audio and video. The probe step records the length, streams, and tags such as title, artist, and album. For anything with a sound track, one ffmpeg pass records the signal step: integrated loudness and loudness range in LUFS, true peak, the silent spans, and the median spectral centroid and flatness. The summary prompt describes these as dynamics, brightness, and texture, and says when a recording is silent throughout. When a resource serves the transcription role, the transcript step sends the first hour to it as 16 kHz mono audio. Each line of the transcript opens with the moment it was said, such as [00:03:12]. When a resource serves the vision role, the scenes step captions a video’s frames: one for each eight seconds, up to eight, taken from the middle of each stretch. Each caption is kept with its moment, and the summary reads them after the transcript.

Every analyzer inherits the same sequence from Analyzer::Base:

  1. Extract children, if the analyzer has any and the feed is less than four levels deep. For example, an email’s attachments become child feeds in the internal children store, and each is analyzed on its own. The analyzer stops here until every child has been analyzed. When the last child finishes, xixo analyzes the parent again.
  2. Derive a thumbnail and a hi-res image for images, video, audio, PDFs, and pages. Audio gets a picture of its waveform. They are stored as references in the internal derived store, under the step derived. Two settings shared by everyone in the tenant size them. Thumbnail width is 320 pixels by default, and a page capture’s thumbnail is the top of the page, cut square. Hi-res size bounds the longest edge, 1,500 pixels by default, and a page capture keeps its full length at that width. A PDF’s first page is drawn at both sizes, so it stays sharp at any setting. A photo or a video frame is never made larger than it is. The hi-res image is what a thumbnail opens and what the vision model reads, except for a GIF: its item page plays the GIF itself when it is 8 MB or less, and a click opens the original. Changing either setting renders an item’s images again the next time it is analyzed.
  3. Describe the file with exiftool under the step metadata: EXIF, XMP, and IPTC from photos, camera and lens details, QuickTime and ID3 tags, PDF and Office document properties, and the like. The step keeps up to 200 named values and drops file system details, maker notes, binary blobs, and container plumbing such as byte offsets and header versions. It skips pages, entries, plain text, JSON, and XML, whose content is their own metadata. exiftool stops after 30 seconds.
  4. Analyze with the analyzer’s own steps, such as text, info, or listing. The download made for the metadata is reused here, so the file is fetched once.
  5. Summarize with the inference resource that serves the analyzer’s role. A text longer than 10,000 characters is summarized in one call from its first 10,000 characters, which open with the names of its sections when it has an outline. The role is smart by default, vision for images and pages, and fast for email and the fallback. The call is sent with the backend’s Effort on routine work as reasoning_effort, so a reasoning model set to none summarizes without thinking first. The prompt fences the file’s text and tells the model to treat it as data. It asks for tags the way a person would label a folder: a topic, the kind of thing the file is, or the name of a person, company, product, or place, and never an amount, a date, an address, or an account or reference number. xixo reduces the model’s answer to a summary, up to 40 entities, and up to 8 tags. It drops stopwords and tags longer than four words. When the analysis finishes, the feed is filed under the tags. If no resource serves the role, there is no summary and no tags.

If derive, analyze, or summarize fails, xixo skips that part and continues. The reference is marked analyzed either way.

AnalyzeFeedJob runs the analyzer. It checks placement first, so a file uploaded to a path it already occupies goes back to the same place. Then it runs the analyzer, unless the feed is an address, and then three more steps:

  • Filing. xixo connects the feed to the xixo:mime feed for its content type.
  • Considering. For an address, and for a file that is still staged, the agent runs under the feed’s own grant (xixo:catalog:read, xixo:catalog:write, xixo:web:read, and xixo:resources:read) with the search, feed, connect, and resource tools. For an address with a schedule, the prompt is the schedule’s prompt. For a staged file, it is FILE_PROMPT, which asks the agent to connect the feed to the tags and feeds it belongs with and lists the places that would accept it. A file synced from a resource is filed by its summary’s tags alone, with no agent. The agent gets the schedule’s turn allowance, or six turns, and stops early after three failed tool calls in a row. If no resource serves the agent role, xixo skips this step.
  • Settling. A file the agent did not place goes to the default storage.

An ask analysis skips all of this. It runs no analyzer, no filing, and no placement. The job answers the question instead, and then rolls the conversation up into the note. Asking describes each part.

An analysis runs for its feed’s timeout, counted from when it starts. setFeedTimeout, or the feed form, sets any value from one minute to one day. A feed with no timeout set runs for ten minutes. An ask on such a feed starts small instead: the time per ask of the model backend that serves the agent role, or two minutes when that is not set. An agent that looks a question up on the web asks for more time while it runs, up to one day from the start. The analysis is done once the answer is saved, and the title and the conversation’s summary follow it. ExpireOverdueJob cancels an analysis that runs past its deadline.

Analysis jobs run on the analysis queue, with at most ANALYSIS_PER_TENANT (2 by default) running per tenant. A Resource::Failed error is retried five times with increasing waits, and the analysis fails if the last retry fails too. An Analyzer::Failed error, or a file with nowhere to go, is discarded and fails the analysis. Any other error fails the analysis and is raised again.

One question is still open. Connecting two feeds does not analyze the feed on the other side again. That needs a way to stop one change from triggering a chain of analyses, such as a cooldown, a depth limit, or a cause that cannot write edges.

Loading a model is slow on a single machine, and two models that do not fit in memory together unload each other on every call. xixo orders analysis work so that each model is used for as long as there is work for it.

An analysis caused by sync reads and files a file in one job: it runs the analyzer, files the feed under its content type and tags, and settles it into storage. An address, or a file that is still staged, needs the agent, so its reading job puts the analysis back in queued and hands the feed to a filing job, which runs the agent. The analysis gets its timeout again when filing starts.

Each job’s queue priority comes from the model its work needs. Lane numbers every model the tenant’s shared inference resources name, starting at 10, so roles served by one model share a priority. The model that serves the agent role always gets the last number. The reading job takes the priority of the analyzer’s summary role, and the filing job takes the priority of the agent role. Solid Queue runs the lowest priority first, so a large sync summarizes everything with one model, then with the next, and files everything with the agent’s model last.

Every other cause runs in a single job at priority 0, ahead of any bulk work, because someone is waiting for it.