Skip to content

Agents and MCP

/mcp is a Model Context Protocol endpoint over streamable HTTP. It accepts the same masks access token as the browser app, as described in tenants. xixo answers a request without a token with 401 and a WWW-Authenticate challenge that names /.well-known/oauth-protected-resource. That document lists the resource URL, the masks issuer as its authorization server, and every xixo: scope an MCP tool checks, with a description. The xixo:settings: scopes are left out, since no tool checks them, and an authorization server that bounds what registered clients may request refuses a request that asks for them. An OAuth client uses the document to find masks and request a token. See connecting an agent.

Calls are rate-limited per token, or per IP address when there is no token, to XIXO_MCP_LIMIT a minute (120 by default). An Mcp-Session-Id belongs to the subject that opened it, and xixo refuses it from anyone else.

Tool Scope to see it Purpose
search xixo:catalog:read searches the whole catalog, or lists the newest feeds of a type
feed xixo:catalog:read reads one feed in full. With do, it adds a note, renames, analyzes, creates, places, or sets how long a feed lasts.
connect xixo:catalog:write connects two feeds, or disconnects them
resource xixo:resources:read lists the resources, and sends a command to one

The MCP reference lists each tool’s inputs.

A token sees only the tools its scopes permit. Tool::Resources.for builds the resource tool for each token. list, describe, runs, get, parameters, and search are always in its schema. check, sync, keep, export, cancel, put, and snapshot appear only when the token has xixo:resources:command. A token with xixo:web:keep also gets snapshot, and can send it only to a resource that renders pages, which is the web type. A command the token cannot use is left out of tools/list.

xixo still checks each call:

  • the feed tool’s writing verbs (note, rename, analyze, create, place, and last) need xixo:catalog:write;
  • search on a resource that serves search needs xixo:web:read;
  • get on a resource that serves fetch, such as curl, needs xixo:web:read.

A token with xixo:mcp:call also sees the tools of every MCP server attached as a resource, passed through unchanged.

Agent runs a prompt against an inference resource that serves the agent role, and offers it the same four tools. It runs for a fixed number of turns: six by default, or the number a feed’s schedule sets, up to 32. Each turn asks the model for a reply.

  • A reply with no tool calls is the answer.
  • xixo checks each tool call against the tool’s input schema before running it. An unknown tool, arguments that are not valid JSON, and arguments the schema refuses go back to the model as an error it can correct. They do not end the loop.
  • A tool result longer than 8,000 characters is truncated, with a note that says how much was cut.
  • Three failed calls in a row end the run as flailed. A successful call resets the count.
  • The last turn offers no tools and tells the model to answer from what it has already read.
  • A run whose analysis is cancelled or past its deadline stops at the start of its next turn, as halted.
  • A run with an analysis is told how much time it has, and is offered more_time. The loop answers more_time itself. Given a number of minutes and a reason, it moves the deadline later, up to one day from the start. The clock tells the model to answer as soon as it has what the request needs and to ask early when it needs longer. When less than a minute is left, or a fifth of the time it was given if that is shorter, the next turn asks for an answer now and offers only more_time.

The run returns an Answer with the reply, the reason the run stopped, the turns taken, and every call made.

When an address is analyzed, the agent runs the schedule’s prompt. When a file that is still staged is analyzed, the agent is asked to file it: it connects the feed to tags and related feeds, and chooses a resource for it. A file synced from a resource is filed without the agent. See analysis.

This agent does not act for a person, so it does not use a person’s token. Feed#grant gives it a grant of its own, with the subject feed: followed by the feed’s key and the scopes in Feed::AGENT_SCOPES:

xixo:catalog:read xixo:catalog:write xixo:web:read xixo:resources:read

The agent can search, read, file, and place feeds, and send read-only commands to resources. It cannot sync, export, write into a resource, call an attached MCP server, or change settings.

Every analysis records requested_by, the subject of the person whose request started it. That is runFeed, analyzeFeed, askCatalog, an upload, or the feed tool called with a person’s token. Analysis#grant builds the agent’s grant from it, and the agent reaches the resources everyone here can use and that person’s own personal resources.

An analysis started by a sync, a schedule, an edge, or another agent records nobody, and its agent reaches only the resources everyone can use. Starting a run on somebody else’s feed reaches your own resources and never theirs.

The grant’s subject stays feed: and the key either way, so the audit trail and the run budget still name the feed. What the agent writes is cataloged for the whole tenant, including anything it read from a personal resource.

xixo saves a question asked of the catalog as a note, and answers it with an analysis whose cause is ask. The note’s key is the question, and the note starts with no title. The analysis is done as soon as the answer is saved, so the answer shows at once. Then the model that answered titles the note from the question and its answer, such as “Toronto today: 18°C and rain”, and rolls the conversation up into the note, with the same Effort on questions. Keeping the whole ask on one model means a machine that holds one large model at a time never swaps models between questions. When no resource serves the agent role, the fast or smart model titles the note. See asking a question.

Answering answers the question in four parts. Only the second and fourth call a model.

  1. Evidence. Evidence chooses what the model reads. It searches the catalog with the question’s keywords, any one of which is enough, fused with the passages and feeds nearest the question in meaning. Up to five feeds are read, in this order: the item the question is about, the feeds earlier answers in the conversation drew on, the two feeds the keywords match best, and then what the fused search found. A note that was itself asked is never read, so an earlier answer is never taken for a source. A feed’s text up to 10,000 characters is read whole, section by section when it has an outline. A longer one is read by up to three sections of its outline, ranked by its passages nearest the question and then by its words, each whole up to 7,000 characters, such as one sheet of a workbook or one minute of a transcript. A table of more than 30 rows is read as its first lines and up to 15 rows that mention the question’s keywords, with a line saying how many rows it has. A feed’s note, and what a model described it as, come first. The evidence stops at 24,000 characters. Up to 20 more of the feeds the search found follow, each as one line with its summary, so a question about which items or how many can see past the five read in full. The model is told that a feed read by several sections is still one item, and that a feed known only by its summary is enough to name but not to quote.
  2. Answer. The model that serves the agent role reads a count of the whole catalog, by type and by kind of file such as image or audio, so it can say how many photos there are, then the evidence, the conversation so far, today’s date with this and last month, quarter, and year spelled out as dates, and the question in one call, and answers in JSON with the feeds it used cited as [feed 12], in the form the question asks for. It is told to copy every number, name, and date as the evidence gives them. The call is sent with the backend’s Effort on questions as reasoning_effort, which is none unless it is set, so a reasoning model answers without thinking first, and at temperature 0, so the same evidence gives the same answer. On qwen3:8b with a sheet of evidence, that answers in about 3 seconds where thinking took 9 to 20.
  3. Compute. When the evidence includes a table, the model asks for a sum, count, average, smallest, or largest over its rows, narrowed by tests on its columns, in place of doing the arithmetic itself. Each table is listed with its columns and two example rows. Tables works it out over the rows the analyzer kept, leaving out a row that is already a total. A column whose numbers are all zero or less, as a ledger writes money spent, is compared by size, so more than 100 means more than 100 either way. The result goes back to the model with how many rows matched, the values they hold most often, and each number column’s range, over the whole table when nothing matched. The model may ask once more with better tests before it answers.
  4. Check. An answer that leaves out the figure a computation gave it is sent back once with the figure named. Then xixo lists every number in the answer that appears nowhere in the evidence, the question, earlier answers, or a computed result, and sends the answer back once with them named. A time of day and a citation are left out of the check.

When the evidence does not answer a question about the world as it is now, such as the weather or a price, the model says so, and xixo hands the question to the agent with the tenant’s web resources: search, fetch, weather, and places. Without any of them, the question is answered from the catalog alone.

Everything runs under the question’s own grant, with the scopes in Feed::ASKING_SCOPES:

xixo:catalog:read xixo:catalog:write xixo:web:read xixo:web:keep xixo:resources:read

A note that was asked holds a conversation. askCatalog with the note’s feedId asks a follow-up question in it, once the previous question has been answered. xixo saves each question on the analysis that answers it, and Ask again asks the latest question again. A follow-up is given the earlier questions and their answers, up to eight of them, with each answer truncated to 1,500 characters, and is searched with the question before it. A follow-up about an earlier answer, such as a request to make it shorter or turn it into a table, is answered by saying that answer again as asked. When the reply is saved, Analyzer::Conversation rolls the whole conversation into the analysis. Every question and answer becomes its text, and the summary, entities, and tags are made from all of it. Search then finds the note by the whole conversation.

Each answer records the feeds it cited in a drew_on step, or the first feed it read when it cites none, and the note’s page lists them under that answer. The note is also connected to all of them.

The agent’s run is confined. Current.confined_to holds the feeds the run created, and the feed tool’s writing verbs refuse any other feed. A page that tells the agent to rewrite a note or rename a file therefore has no effect. The run can create only notes, never a scheduled feed. connect needs one end to be the question or a feed the run created. Snapshots are not confined, because the page decides what a snapshot contains.

A feedback resource answers nothing. Its one command, ask, takes a question and, optionally, the context it came up in and what was wanted. Every ask returns answered: false and the same sentence telling the caller to carry on, and is kept as a note titled “Wanted:” and the question. The note holds the question, its context, what would have helped, the grant’s subject, and the feed the caller was working on, and it is connected to that feed and to the /feedback address.

Attaching a feedback resource makes /feedback, so it sits on the shelf beside Everything and lists everything anyone wanted and could not get. Its schedule runs once a week with 16 turns: its agent reads what was wanted, groups it by what would provide it, such as a resource to attach or a change to xixo, and keeps a note titled “Feedback review” with each group’s count, a few of its questions, and the change it proposes. A health check makes /feedback again if it is gone.

The resource tool tells every MCP client to use it for anything no tool could answer or do, and a tenant’s web agent is told the same. When an answer from the catalog finds nothing, the model says what would have answered the question, and xixo asks the tenant’s feedback resource with it. A tenant without a feedback resource keeps nothing.

Work that can run for a long time spends from a budget. This covers feed with do: analyze, and resource with sync or export. Each subject in a tenant may start XIXO_RUN_BUDGET of these per clock hour, 20 by default, and 0 removes the limit. A feed’s agent spends from its own subject’s budget, which is separate from any person’s.

A gate is a row that stops a kind of job. The iterating jobs each name their gate: sync, analyze, export, or reindex. A gate applies to the whole key or to one record, such as a single resource for sync, and the record’s gate takes precedence over the key’s. A disabled gate marks the run as gated and does nothing. A gate that is enabled but not live lets sync and reindex walk and count without writing anything. Jobs read their gates again every 50 iterations, so closing a gate stops a run that is already in progress. XIXO_ITERATORS_DISABLED closes every gate at once. Gates have no screen yet. Set them with Gate.set! from a console.

Every tool call writes an audit event with the tool, the scope, the result (ok, error, or denied), the IP address and request id, and the duration. Each event names its actor: a person with their name and the client they asked through, a client acting for itself, the agent, or nobody when no token was accepted. An agent’s event also names the analysis it ran for, so the Activity page in Settings shows one analysis as one entry. Each event carries a sentence about the call, such as “filed invoice.pdf under receipts”, and the feed it was about. xixo writes the sentence when the call is made, so it still reads correctly after the feed is forgotten. xixo summarizes the arguments and does not store them in full. It truncates strings to 200 characters, counts arrays, and keeps one level of nesting. It replaces the value of any key that names a secret, password, token, credential, authorization, or access key with [redacted]. Refused authorizations, rate-limited calls, and some GraphQL mutations are recorded the same way.

The auditEvents query returns events to a token with xixo:catalog:read. A daily job deletes events older than XIXO_AUDIT_RETENTION_DAYS, 90 by default, and 0 keeps them forever.