Talk to the LLM server via GET /v1/models and POST /v1/chat/completions
(SSE) instead of Ollama's native API. Any OpenAI compatible server now
works (Ollama, vLLM, llama.cpp, LM Studio, Hugging Face TGI, ...);
Ollama serves this API natively, existing setups keep working unchanged.
The Ollama specific keep_alive option is dropped, the ai_summary.grounding
setting is added as instance wide default of the grounding preference.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add an optional, disabled-by-default plugin that shows an AI generated
summary at the top of the result page, generated by a (local) Ollama
server:
- async: the result page is never delayed; a client plugin streams the
answer (NDJSON over a new /ai_summary endpoint) into a placeholder
answer with a typing indicator, collapsed behind a More button, with
an inline follow-up chat
- trigger: first page of general searches only, skipped when an infobox
or instant answer already answers the query
- grounding (per-user preference): send the top result snippets as
context, the model answers from them instead of its own knowledge
- configuration: new AI Summary preferences tab (server URL, model,
grounding) with instance defaults in a new ai_summary: settings
section; all three preferences can be locked for public instances
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>