text-summarize/MODULE.de.md
flemming-it 86f191ca3e
Some checks failed
CI / Linux x86_64 (Forgejo) (push) Failing after 2s
sign-bundle / sign (push) Failing after 1m45s
feat: api accepts 'vllm' as a first-class value
'openai' reads as the cloud company — an operator running
self-hosted vLLM should be able to write what they mean. The new
value speaks the identical OpenAI wire (a guard test asserts the
emitted request stays byte-identical to api: openai, so the alias
can never drift into a dialect) but yields vLLM-specific error
hints (vllm serve, port 8000) instead of generic OpenAI prose.
Manifest + docs name the value in DE and EN.

Signed-off-by: flemming-it <sf@flemming.it>
2026-08-20 17:13:43 +02:00

3.8 KiB

text.summarize

LLM-basierte treue Zusammenfassung (Ollama, vLLM, jeder OpenAI-kompatible Server oder Anthropic). Sendet Quelltext an einen konfigurierten LLM-Endpunkt mit einem Treue-vor-Kreativität- System-Prompt und liefert die Zusammenfassung plus audit-taugliche Modell-Herkunfts-Felder.

Capability

  • text.summarize@0.1.0

Eingaben

Name Typ Beschreibung
text text Quelltext.
style text Optionaler Stil-Hinweis (z.B. one paragraph, three bullet points, a tweet). Default: one paragraph.
language text Optionaler Ziel-Sprache (z.B. German, ja-JP). Leer = Quellsprache.
endpoint text Chat-Endpunkt passend zum gewählten api (Ollama /api/chat, OpenAI-kompatibel /v1/chat/completions — vLLM u. a., Anthropic /v1/messages).
api text Optionales Wire-Format: ollama (Default), vllm, openai, anthropic (vllm = OpenAI-Wire mit vLLM-spezifischen Hinweisen).
model text Modell-ID am Endpunkt.
api_key text Optionaler Bearer-Token für Cloud-Endpunkte.

Ausgaben

Name Typ Beschreibung
summary text Die Zusammenfassung.
style text Echo des Stil-Inputs.
language text Echo des Sprach-Inputs.
model_endpoint text URL, gegen die zusammengefasst wurde.
model_name text Modell-ID wie an die LLM-API gesendet.
model_digest text SHA-256-Digest des bedienten Modells — nur bei Ollama, sonst leer.

Treue vor Kreativität

Der System-Prompt ist so getrimmt, dass die Zusammenfassung aus dem Quelltext herleitbar bleibt — keine Spekulation, kein mitgedachter Kontext, keine „kreativen" Umformulierungen, die abdriften. Geeignet für Compliance-Flows, in denen eine halluzinierte Zusammenfassung schlimmer ist als eine ausführliche. Das Modell behält die Freiheit, die das vom Operator gewählte LLM ihm gibt; dieses Modul ist der Prompt

  • Audit-Rahmen, kein fine-tuned Modell.

Berechtigungen

permissions:
  - "net: localhost"
  - "net: 127.0.0.1"
  - "net: api.openai.com"
  - "net: api.anthropic.com"

Analog text.translate — lokales Ollama per Default, Cloud per Operator-Policy zuschaltbar.

Grenzen in v0.1.0

  • Single-Shot. Sehr lange Quelltexte (> 32k Tokens je nach Modell) brauchen einen vorgeschalteten Chunker; v0.1.0 macht noch keine automatische Chunked-Summarise+Reduce- Pipeline.
  • Keine strukturierte Ausgabe. summary ist Freitext. JSON-formatierte Zusammenfassungen ({ key_points: [], conclusions: [] }) sind ein v0.2.0-Kandidat.

Beispiel-Flow

name: extract-summarise
inputs:
  document: bytes
steps:
  - id: extract
    use: text.extract@^0
    with:
      document: $inputs.document
  - id: summarise
    use: text.summarize@^0
    with:
      text: $extract.extracted.pages[*].text
      style: "three bullet points"
      endpoint: "http://localhost:11434/api/chat"
      model: "qwen2.5:14b"
outputs:
  summary: $summarise.summary
  audit_model: $summarise.model_digest

Build

cargo build --release --target wasm32-wasip2
# Ausgabe: target/wasm32-wasip2/release/text_summarize.wasm