feat: law-source-bund 0.1.0 (source.bundesrecht)
All checks were successful
CI / Linux x86_64 (Forgejo) (push) Successful in 1m42s

Fetch one German federal law from gesetze-im-internet.de (BMJ) by its
site slug, unpack the BMJ norm XML from the site's xml.zip and return
it content-addressed (SHA-256) plus the BMJ builddate, so a flow's
audit trail pins the exact Gesetzesstand it processed.

- xml.zip unpack picks the largest XML entry (norm body)
- lying ZIP size header rejected instead of silently truncated
  (audit integrity); 64 MB unpack cap, 96 MB download cap
- net permission pinned to www.gesetze-im-internet.de; no auth
- 11 unit tests (no network), wasm32-wasip2 build green

Signed-off-by: flemming-it <sf@flemming.it>
This commit is contained in:
flemming-it 2026-07-17 00:29:06 +02:00
commit f1652cce62
16 changed files with 1760 additions and 0 deletions

59
MODULE.md Normal file
View file

@ -0,0 +1,59 @@
# source.bundesrecht
<!-- chain:io-card:start -->
<!-- Generated by `chain doc` from module.yaml — do not edit by hand. -->
![law-source-bund inputs and outputs](law-source-bund.io.svg)
<!-- chain:io-card:end -->
Fetch German federal law from gesetze-im-internet.de by slug
## What it does
Resolves a gesetze-im-internet.de slug (e.g. `bgb`, `estg`, `gg`,
`stromnzv`) against the public BMJ publication site, downloads the
law's `xml.zip`, unpacks the BMJ norm XML (gii-norm DTD) and returns
it plus a metadata record. The unpacked bytes are content-addressed
(SHA-256), so the flow's audit trail pins the exact Gesetzesstand
that was processed.
Unlike EUR-Lex, gesetze-im-internet serves only the CURRENT
consolidated state — there is no URL for an older version. The
SHA-256 (plus the BMJ `builddate` from the XML header) is what makes
a later re-fetch comparable: a change upstream becomes visible
instead of silent.
Public BMJ data only; network permission is pinned to
`www.gesetze-im-internet.de`.
## How to use
```yaml
steps:
- id: fetch
use: source.bundesrecht@^0
with:
gesetz: stromnzv
- id: normalize
use: text.akoma-normalize@^0
with:
content: $fetch.xml
```
The slug is the path segment of the law's page on
gesetze-im-internet.de: `https://www.gesetze-im-internet.de/<slug>/`.
## Outputs
- `xml` — the BMJ norm XML as unpacked from the site's archive.
Feed into `text.akoma-normalize` for the structured
Akoma-Ntoso-aligned representation (its v0.1 BMJ adapter handles
this format natively).
- `meta` — JSON SourceMeta: `gesetz`, `url`, `source_sha256`,
`byte_len`, and `builddate` when the XML header carries one.
## Errors
- Unknown slug → clear invalid-input error naming the checked URL.
- Non-ZIP or XML-less download, HTML error pages → internal error
with the reason; nothing HTML-shaped is ever passed downstream.