[doc] ai_summary: rewrite for readers new to the feature
The documentation was written from the perspective of someone who had just implemented the plugin: it opened with SSRF caveats and settings keys, and explained decisions rather than usage. The admin page now starts from what the feature is, followed by a four-step quickstart (install Ollama, pull a model, configure, restart) and a troubleshooting section for the failures that actually occur. The quickstart repeats the default plugins because a plugins: block replaces that list instead of merging into it -- following the short version of the instructions would otherwise switch every other plugin off. The developer page gains a rendered data flow diagram. The two-request design -- placeholder first, streamed answer second -- is the part of this plugin that is hard to convey in prose, and the SSE to NDJSON change is easier to see than to read about. The module docstring now describes the module and links to both pages, instead of restating administration guidance. Add Jason Witty to AUTHORS.rst. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
8ff85790cf
commit
b6eed9a993
+13
-55
@@ -1,64 +1,22 @@
|
||||
# SPDX-License-Identifier: AGPL-3.0-or-later
|
||||
"""Plugin that displays an AI generated summary of the search query at the top
|
||||
of the result page. The summary is generated by a (local) LLM server that
|
||||
implements the `OpenAI chat completions API`_ -- e.g. `Ollama`_, vLLM,
|
||||
llama.cpp, LM Studio or Hugging Face TGI.
|
||||
"""Implementation of the AI summary plugin, which shows a generated answer above
|
||||
the search results. The answer comes from an LLM server that implements the
|
||||
`OpenAI chat completions API`_ (Ollama, vLLM, llama.cpp, LM Studio, Hugging Face
|
||||
TGI, ...) and that the administrator runs.
|
||||
|
||||
The LLM server URL and the model are configured by the user in the *AI
|
||||
Summary* tab of the preferences (``ai_summary_server``, ``ai_summary_model``);
|
||||
the administrator can configure instance wide defaults in the ``ai_summary:``
|
||||
section and lock the preferences via :ref:`settings preferences`.
|
||||
- :ref:`ai_summary plugin` describes the design and the request flow.
|
||||
- :ref:`settings ai_summary` describes how to configure it.
|
||||
|
||||
.. attention::
|
||||
This module holds the plugin itself and the ``/ai_summary`` endpoint
|
||||
(:py:obj:`ai_summary_view`, registered in :py:obj:`searx.webapp`). The endpoint
|
||||
streams the answer to the browser, so that the result page is never delayed by
|
||||
the LLM; :py:obj:`SXNGPlugin.post_search` only adds an empty
|
||||
:py:obj:`searx.result_types.AiSummary` placeholder for the client to fill.
|
||||
|
||||
A user configurable server URL allows any user of the instance to make the
|
||||
SearXNG server send requests to a URL of their choice (`SSRF`_), and each
|
||||
summary is real LLM work. This plugin is intended for private instances --
|
||||
on a public instance, lock the ``ai_summary_server``, ``ai_summary_model``
|
||||
and ``ai_summary_grounding`` preferences and configure the ``ai_summary:``
|
||||
section instead.
|
||||
Settings of the ``ai_summary:`` section are defined in
|
||||
:py:obj:`searx.ai_summary.SettingsAISummary`.
|
||||
|
||||
The result page is never delayed by this plugin: it only places an empty
|
||||
placeholder (:py:obj:`searx.result_types.AiSummary`) in the answer area, which
|
||||
is filled asynchronously by the client (``client/simple/src/js/plugin/
|
||||
AiSummary.ts``) from the ``/ai_summary`` endpoint (registered in
|
||||
:py:obj:`searx.webapp`). The endpoint re-emits the SSE token stream of the
|
||||
LLM server's ``/v1/chat/completions`` to the client as `NDJSON`_.
|
||||
|
||||
A summary is only generated on the first page of a *general* search and only
|
||||
if no engine has contributed an infobox (e.g. wikipedia / wikidata) or an
|
||||
instant answer (e.g. ddg definitions) -- in these cases the query is most
|
||||
likely a lookup of a well known term that is already answered.
|
||||
|
||||
The requests to the LLM server are sent directly (not via
|
||||
:py:obj:`searx.network`), an outgoing proxy configuration is deliberately not
|
||||
applied to reach an LLM server in the local network.
|
||||
|
||||
A server that requires authentication (e.g. vLLM or llama.cpp started with
|
||||
``--api-key``, or an LLM server behind an authenticating reverse proxy) is
|
||||
configured with an ``api_key``. The administrator's key is only sent to the
|
||||
server in ``base_url``, never to a server a user configured; for their own
|
||||
server users configure their own key in the ``ai_summary_api_key`` preference
|
||||
(:py:obj:`_server_api_key`).
|
||||
|
||||
Configuration of the defaults (:py:obj:`searx.ai_summary.SettingsAISummary`):
|
||||
|
||||
.. code:: yaml
|
||||
|
||||
ai_summary:
|
||||
base_url: "http://127.0.0.1:11434"
|
||||
model: "llama3.2:3b"
|
||||
|
||||
.. code:: yaml
|
||||
|
||||
plugins:
|
||||
searx.plugins.ai_summary.SXNGPlugin:
|
||||
active: false
|
||||
|
||||
.. _Ollama: https://ollama.com/
|
||||
.. _OpenAI chat completions API: https://platform.openai.com/docs/api-reference/chat
|
||||
.. _NDJSON: https://github.com/ndjson/ndjson-spec
|
||||
.. _SSRF: https://owasp.org/www-community/attacks/Server_Side_Request_Forgery
|
||||
"""
|
||||
|
||||
import typing as t
|
||||
|
||||
Reference in New Issue
Block a user