[doc] ai_summary: rewrite for readers new to the feature

The documentation was written from the perspective of someone who had
just implemented the plugin: it opened with SSRF caveats and settings
keys, and explained decisions rather than usage.

The admin page now starts from what the feature is, followed by a
four-step quickstart (install Ollama, pull a model, configure, restart)
and a troubleshooting section for the failures that actually occur.  The
quickstart repeats the default plugins because a plugins: block replaces
that list instead of merging into it -- following the short version of
the instructions would otherwise switch every other plugin off.

The developer page gains a rendered data flow diagram.  The two-request
design -- placeholder first, streamed answer second -- is the part of
this plugin that is hard to convey in prose, and the SSE to NDJSON
change is easier to see than to read about.

The module docstring now describes the module and links to both pages,
instead of restating administration guidance.

Add Jason Witty to AUTHORS.rst.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
jasonwitty
2026-08-10 12:30:53 -07:00
co-authored by Claude Opus 5
parent 8ff85790cf
commit b6eed9a993
4 changed files with 326 additions and 103 deletions
+13 -55
View File
@@ -1,64 +1,22 @@
# SPDX-License-Identifier: AGPL-3.0-or-later
"""Plugin that displays an AI generated summary of the search query at the top
of the result page. The summary is generated by a (local) LLM server that
implements the `OpenAI chat completions API`_ -- e.g. `Ollama`_, vLLM,
llama.cpp, LM Studio or Hugging Face TGI.
"""Implementation of the AI summary plugin, which shows a generated answer above
the search results. The answer comes from an LLM server that implements the
`OpenAI chat completions API`_ (Ollama, vLLM, llama.cpp, LM Studio, Hugging Face
TGI, ...) and that the administrator runs.
The LLM server URL and the model are configured by the user in the *AI
Summary* tab of the preferences (``ai_summary_server``, ``ai_summary_model``);
the administrator can configure instance wide defaults in the ``ai_summary:``
section and lock the preferences via :ref:`settings preferences`.
- :ref:`ai_summary plugin` describes the design and the request flow.
- :ref:`settings ai_summary` describes how to configure it.
.. attention::
This module holds the plugin itself and the ``/ai_summary`` endpoint
(:py:obj:`ai_summary_view`, registered in :py:obj:`searx.webapp`). The endpoint
streams the answer to the browser, so that the result page is never delayed by
the LLM; :py:obj:`SXNGPlugin.post_search` only adds an empty
:py:obj:`searx.result_types.AiSummary` placeholder for the client to fill.
A user configurable server URL allows any user of the instance to make the
SearXNG server send requests to a URL of their choice (`SSRF`_), and each
summary is real LLM work. This plugin is intended for private instances --
on a public instance, lock the ``ai_summary_server``, ``ai_summary_model``
and ``ai_summary_grounding`` preferences and configure the ``ai_summary:``
section instead.
Settings of the ``ai_summary:`` section are defined in
:py:obj:`searx.ai_summary.SettingsAISummary`.
The result page is never delayed by this plugin: it only places an empty
placeholder (:py:obj:`searx.result_types.AiSummary`) in the answer area, which
is filled asynchronously by the client (``client/simple/src/js/plugin/
AiSummary.ts``) from the ``/ai_summary`` endpoint (registered in
:py:obj:`searx.webapp`). The endpoint re-emits the SSE token stream of the
LLM server's ``/v1/chat/completions`` to the client as `NDJSON`_.
A summary is only generated on the first page of a *general* search and only
if no engine has contributed an infobox (e.g. wikipedia / wikidata) or an
instant answer (e.g. ddg definitions) -- in these cases the query is most
likely a lookup of a well known term that is already answered.
The requests to the LLM server are sent directly (not via
:py:obj:`searx.network`), an outgoing proxy configuration is deliberately not
applied to reach an LLM server in the local network.
A server that requires authentication (e.g. vLLM or llama.cpp started with
``--api-key``, or an LLM server behind an authenticating reverse proxy) is
configured with an ``api_key``. The administrator's key is only sent to the
server in ``base_url``, never to a server a user configured; for their own
server users configure their own key in the ``ai_summary_api_key`` preference
(:py:obj:`_server_api_key`).
Configuration of the defaults (:py:obj:`searx.ai_summary.SettingsAISummary`):
.. code:: yaml
ai_summary:
base_url: "http://127.0.0.1:11434"
model: "llama3.2:3b"
.. code:: yaml
plugins:
searx.plugins.ai_summary.SXNGPlugin:
active: false
.. _Ollama: https://ollama.com/
.. _OpenAI chat completions API: https://platform.openai.com/docs/api-reference/chat
.. _NDJSON: https://github.com/ndjson/ndjson-spec
.. _SSRF: https://owasp.org/www-community/attacks/Server_Side_Request_Forgery
"""
import typing as t