New optional input api: ollama|openai|anthropic (default ollama — existing flows unchanged). - openai: OpenAI Chat Completions format (/v1/chat/completions), Bearer auth — targets vLLM and compatible self-hosted servers. - anthropic: Messages API (/v1/messages), x-api-key + anthropic-version headers, top-level system field, mandatory max_tokens (fixed 4096). - Audit outputs unchanged: model_endpoint/model_name always set; model_digest stays Ollama-only (no digest API on openai/anthropic, field is left empty rather than fabricated). - api_key is only ever placed in auth headers; never in outputs, errors, or audit events (verified against the event log). - No streaming, no tool calls. Bump module + capability version to 0.2.0. Wire-format unit tests for request serialization and response parsing against fixed JSON fixtures; live smoke green on the ollama path and on the openai path against an OpenAI-compatible local endpoint. Signed-off-by: flemming-it <sf@flemming.it>
5.4 KiB
llm.chat
Generic LLM chat adapter. The building-block module any flow
uses when it needs a one-shot LLM completion — classify,
extract-fields, rewrite, decide — instead of rolling its own
HTTP client. Speaks three wire formats, selected by the
optional api input: Ollama (default), OpenAI Chat
Completions (OpenAI, vLLM, compatible servers), and the
Anthropic Messages API.
Capability
llm.chat@0.2.0
Inputs
| Name | Type | Description |
|---|---|---|
prompt |
text | The user-facing prompt. |
endpoint |
text | Chat endpoint URL matching the selected api (see below). |
model |
text | Model identifier (e.g. qwen2.5:14b, claude-fable-5). |
api_key |
text | Optional API key. Never appears in outputs, logs, or events. |
system_prompt |
text | Optional system message. Empty = use the model's default. |
api |
text | Optional wire format: ollama (default), openai, anthropic. |
Wire formats (api)
api |
Endpoint shape | Auth header | Notes |
|---|---|---|---|
ollama |
http://localhost:11434/api/chat |
Authorization: Bearer (optional) |
v0.1.0 behavior, unchanged. Model-digest probe via /api/show. |
openai |
https://.../v1/chat/completions |
Authorization: Bearer |
OpenAI Chat Completions format — also spoken by vLLM and most self-hosted inference servers. |
anthropic |
https://api.anthropic.com/v1/messages |
x-api-key + anthropic-version |
Messages API: system is a top-level field; max_tokens is mandatory and defaults to 4096. |
Outputs
| Name | Type | Description |
|---|---|---|
response |
text | The assistant's plain-text reply. |
model_endpoint |
text | The endpoint the response was generated against. |
model_name |
text | The model identifier as supplied to the LLM API. |
model_digest |
text | SHA-256 digest of the served Ollama model. Empty for openai/anthropic — those APIs expose no digest; the field is left empty rather than fabricated. |
Together the three model_* outputs answer the audit
question: "which exact model produced this response?" — that
audit trail is the reason flows use this module rather than
spawning their own HTTP calls.
Permissions
permissions:
- "net: localhost"
- "net: 127.0.0.1"
- "net: api.openai.com"
- "net: api.anthropic.com"
Loopback (local Ollama / vLLM) by default. Cloud endpoints require an
operator-policy override in ~/.chain/config.yaml#security.max_permissions.
Why this module instead of inline HTTP
Three reasons:
- Audit. Every LLM call surfaces
model_endpoint,model_name,model_digestas separate outputs that land in the hash-chained audit log alongside the response. A regulator-facing reproducibility check just compares the digest field to the model snapshot. - Permission posture. The operator sees
llm.chatin the installed modules + its declarednet:list. A home-grown HTTP module would either hide its endpoints or be a fresh review surface every time. - Endpoint portability. The endpoint, model, and wire
format are flow inputs, not compile-time constants. The
same flow runs against
localhost:11434in dev and a production vLLM cluster in prod just by swapping inputs.
Limits in v0.2.0
- No streaming. The whole reply is buffered before the output step fires.
- No tool-call / function-call surface. A flow needing tool
use composes multiple
llm.chatsteps with prompt engineering, or uses MCP via the bridge. - Anthropic
max_tokensis fixed at 4096 (the API requires the field; a configurable input lands when a flow needs it).
Example flows
name: classify-incoming
inputs:
text: text
steps:
- id: classify
use: llm.chat@^0
with:
prompt: $inputs.text
system_prompt: |
Classify the text as one of: question, complaint,
feedback, spam. Answer with the label only.
endpoint: "http://localhost:11434/api/chat"
model: "qwen2.5:14b"
outputs:
category: $classify.response
audit_model: $classify.model_digest
Against a vLLM (OpenAI-compatible) server:
steps:
- id: classify
use: llm.chat@^0
with:
api: "openai"
endpoint: "http://inference.internal:8000/v1/chat/completions"
model: "meta-llama/Llama-3.1-8B-Instruct"
prompt: $inputs.text
Against the Anthropic Messages API:
steps:
- id: classify
use: llm.chat@^0
with:
api: "anthropic"
endpoint: "https://api.anthropic.com/v1/messages"
model: "claude-fable-5"
api_key: $inputs.anthropic_key
prompt: $inputs.text
Build
cargo build --release --target wasm32-wasip2
# Output: target/wasm32-wasip2/release/llm_chat.wasm