documents.content)
and indexes it in FTS5. The LLM classifier plugin is already
configured for the operator-acknowledged endpoint. Combining the two
into a natural-language Q&A layer is a natural next feature — but a
big one. This doc is the plan.
Status: wishlist. No code. This is what future PRs land against.
Motivating queries
The operator’s own words:- “when is my car insurance due?”
- “what was the total utilities amount paid in month of November last year?”
- “who signed the lease amendment?”
Design
Layer 1: retrieval — FTS5 + optional embeddings
documents.content + FTS5 already answer “which docs contain X.”
That’s enough for the first shape.
For richer semantic recall (“insurance renewal” matches “policy
expires”), pair with sqlite-vec embeddings (already tracked as
#128 sweep-2 C3). Same LLM endpoint provides the vectors — no new
egress surface. When the extension isn’t loaded, degrade to FTS5-
only. Zero-egress installs get less recall but not less privacy.
Layer 2: constrained prompt + JSON extraction
Prompt template (server-side, versioned):- Excerpts, not full text. Snippet extraction reuses the search path’s FTS5 highlight. Keeps prompts small + hides the rest of the archive from the LLM even when the operator’s endpoint is their own local Ollama.
- Cited doc ids. The response must name the docs it read. Answers without citations get discarded. This is the load-bearing hallucination-check.
- Confidence field. Low-confidence answers (”< 0.5”) render as “I found some possibly-related docs; check them yourself” with the citations linked.
Layer 3: aggregation queries via structured extraction
For “total utilities in November last year,” the LLM can’t sum reliably — but it CAN extract{doc_id, amount, currency, date}
structured records per matched doc. Server does the sum. Prompt
shape:
Layer 4: HTTP surface
documents:read. Falls back to a plain search result when
the LLM endpoint isn’t configured. Rate-limited (5rps / burst 10)
per source IP — expensive endpoint, cheap DoS.
Layer 5: MCP tool
Same call surfaces through the MCP server as anask tool. Agents
(Claude Desktop, Cursor, custom) can ask the archive questions
without wiring HTTP.
Wire prerequisites
- #128 (sqlite-vec) landed. Semantic recall is optional but a big win.
- LLM endpoint acknowledged. Same
llm.endpoint_url+llm.egress_ackconfig the classifier uses. - New config:
ASK_MAX_TOKENS(default 4096),ASK_TIMEOUT(default 30s),ASK_MAX_DOCS(default 8, cap 20). Everything goes through the existingsettingstable via the setup wizard.
Privacy invariants
documents.contentnever leaves the server outside the LLM request. The response we cache locally is the LLM’s answer + the ids, not the excerpts we shipped.- The excerpt window is bounded (default 400 chars per doc); we never send the whole document.
- A doc marked
sensitivity IN ('confidential','restricted')is excluded from the retrieval pool by default. An override header or setting can flip that for personal instances; the archive admin explicitly toggles it.
What NOT to build
- Chat history. One question, one answer, no context carry. Multi-turn is a v2; single-shot covers the motivating queries.
- Fine-tuning. The LLM never sees the archive as training data.
- Automatic re-answer on doc mutation. Answers are a query-time computation; caching stale answers to a mutable archive is a correctness footgun.
Rollout order
- Design doc lands ✅ (this file).
- #128 sqlite-vec extension + embedding-at-ingest job.
POST /api/askhandler wiring FTS5 → prompt → LLM → JSON parse → cite-check.- MCP tool alias.
- SPA route (Search screen grows an “Ask” tab).
- Operator docs (how to bring up a local Ollama for offline use).
ASK_ENABLED=false by default) until the citation-check quality
is proven on the maintainer’s own archive.