text-summarize/MODULE.md
flemming-it 86f191ca3e
Some checks failed
CI / Linux x86_64 (Forgejo) (push) Failing after 2s
sign-bundle / sign (push) Failing after 1m45s
feat: api accepts 'vllm' as a first-class value
'openai' reads as the cloud company — an operator running
self-hosted vLLM should be able to write what they mean. The new
value speaks the identical OpenAI wire (a guard test asserts the
emitted request stays byte-identical to api: openai, so the alias
can never drift into a dialect) but yields vLLM-specific error
hints (vllm serve, port 8000) instead of generic OpenAI prose.
Manifest + docs name the value in DE and EN.

Signed-off-by: flemming-it <sf@flemming.it>
2026-08-20 17:13:43 +02:00

3.6 KiB

text.summarize

LLM-backed faithful summarisation. Sends source text to a configured LLM endpoint with a fidelity-over-creativity system prompt and emits the summary plus audit-grade model-provenance fields.

Capability

  • text.summarize@0.1.0

Inputs

Name Type Description
text text Source text to summarise.
style text Optional style hint (e.g. one paragraph, three bullet points, a tweet). Default: one paragraph.
language text Optional output-language hint (e.g. German, ja-JP). Empty = same as source.
endpoint text Chat endpoint URL matching the selected api (Ollama /api/chat, OpenAI-compatible /v1/chat/completions — vLLM etc., Anthropic /v1/messages).
api text Optional wire format: ollama (default), vllm, openai, anthropic (vllm = the OpenAI wire with vLLM-specific hints).
model text Model identifier the endpoint serves.
api_key text Optional bearer token for cloud-hosted endpoints.

Outputs

Name Type Description
summary text The summary.
style text Echo of the input style.
language text Echo of the input language.
model_endpoint text URL the summary was generated against.
model_name text Model identifier as supplied to the LLM API.
model_digest text SHA-256 digest of the served model — Ollama only, empty elsewhere.

Fidelity over creativity

The system prompt is tuned so the summary is derivable from the source — no speculation, no implied context, no "creative" rephrasing that drifts. Suitable for compliance flows where a hallucinated summary is worse than a verbose one. The model still has the freedom the operator's chosen LLM gives it; this module is the prompt + audit harness, not a fine-tuned model.

Permissions

permissions:
  - "net: localhost"
  - "net: 127.0.0.1"
  - "net: api.openai.com"
  - "net: api.anthropic.com"

Same shape as text.translate — local Ollama by default, opt-in cloud via operator policy.

Limits in v0.1.0

  • Single-shot. Very long source texts (>32k tokens depending on the model) need an upstream chunker; v0.1.0 does not yet do automatic chunked-summarise + reduce.
  • No structured output. summary is free-form text. JSON- formatted summaries (e.g. { key_points: [], conclusions: [] }) are a v0.2.0 candidate.

Example flow

name: extract-summarise
inputs:
  document: bytes
steps:
  - id: extract
    use: text.extract@^0
    with:
      document: $inputs.document
  - id: summarise
    use: text.summarize@^0
    with:
      text: $extract.extracted.pages[*].text
      style: "three bullet points"
      endpoint: "http://localhost:11434/api/chat"
      model: "qwen2.5:14b"
outputs:
  summary: $summarise.summary
  audit_model: $summarise.model_digest

Build

cargo build --release --target wasm32-wasip2
# Output: target/wasm32-wasip2/release/text_summarize.wasm