'openai' reads as the cloud company — an operator running
self-hosted vLLM should be able to write what they mean. The new
value speaks the identical OpenAI wire (a guard test asserts the
emitted request stays byte-identical to api: openai, so the alias
can never drift into a dialect) but yields vLLM-specific error
hints (vllm serve, port 8000) instead of generic OpenAI prose.
Manifest + docs name the value in DE and EN.
Signed-off-by: flemming-it <sf@flemming.it>
New optional input api: ollama|openai|anthropic (default ollama —
existing flows unchanged).
- openai: OpenAI Chat Completions format (/v1/chat/completions),
Bearer auth — targets vLLM and compatible self-hosted servers.
- anthropic: Messages API (/v1/messages), x-api-key +
anthropic-version headers, top-level system field, mandatory
max_tokens (fixed 4096).
- Audit outputs unchanged: model_endpoint/model_name always set;
model_digest stays Ollama-only (no digest API on openai/anthropic,
field is left empty rather than fabricated).
- api_key is only ever placed in auth headers; never in outputs,
errors, or audit events (verified against the event log).
- No streaming, no tool calls.
Bump module + capability version to 0.2.0. Wire-format unit tests
for request serialization and response parsing against fixed JSON
fixtures; live smoke green on the ollama path and on the openai
path against an OpenAI-compatible local endpoint.
Signed-off-by: flemming-it <sf@flemming.it>
SDK 0.3.0 drops the fai-legacy #[fai_module] alias; pin the tag and
flip the attribute, rebuilt module.wasm accordingly.
Signed-off-by: flemming-it <sf@flemming.it>
Generic Ollama-compatible LLM chat adapter. The lower-level
counterpart to orchestrator-llm: orchestrator-llm wraps the LLM
in a planning prompt that emits a structured F∆I Plan; this
module is the plain-prompt adapter that flows compose for
summarisation, translation, free-form Q&A, etc.
Capability surface:
Inputs:
prompt : text
endpoint : text (Ollama /api/chat URL)
model : text
api_key : text (optional bearer token)
system_prompt : text (optional)
Outputs:
response : text (assistant reply)
model_endpoint : text (audit correlation)
model_name : text (audit correlation)
model_digest : text (Ollama /api/show probe; empty
for non-Ollama or transient failures)
Permissions: net to localhost / 127.0.0.1 / api.openai.com /
api.anthropic.com.
The audit-field trio (endpoint + name + digest) closes the same
forensic gap that orchestrator-llm v0.3.1 closed for plan
generation: any historical chat invocation can be traced to the
exact model that produced it.
Implementation reuses the same defensive Ollama client pattern
from orchestrator-llm — derive_show_url + extract_show_digest
+ best-effort probe_model_digest. Duplication accepted at the
two-module mark; a shared crate refactor lands once a third
module needs the same plumbing.
12 host-side tests cover prompt building, Ollama-shaped
response parsing, URL transform, digest extraction (top-level
+ nested), end-to-end success, end-to-end probe-failure
swallow, end-to-end skip-for-non-Ollama, and the missing-input
guards.
Wasm artifact: 294 KB. Verified to build with v1.0 fai:platform
imports baked in.
Bootstrapped via 'fai new module llm.chat' (workspace v0.10.13)
which now produces an SDK-based template directly.
Signed-off-by: flemming-it <sf@flemming.it>