aboutsummaryrefslogtreecommitdiff
path: root/src/main/java
Commit message (Collapse)AuthorAgeFilesLines
* Die Startseite führt die Neuankömmlinge einMatthias Andreas Benkard9 days5-11/+297
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Seit der Inbetriebnahme setzte die Startseite voraus, was gerade nicht hat, wer sie zum ersten Mal aufruft: geeignete Dateien. Über dem Formular stand die Tätigkeitsbeschreibung, darunter zwei leere Dateifelder, und wer nie mit gii-Norm-XML oder einem BGBl-Regelungstext zu tun hatte, konnte das Werkzeug nicht einmal ausprobieren. Ein zugeklappter Block tritt hinzu, der zwei durchgerechnete Fälle mit Verweisen auf die amtlichen Fundstellen anbietet. Zugeklappt kostet er eine Zeile; das Formular bleibt, wo es war. Angeboten werden ein verkündetes Änderungsgesetz — das UWG nebst dem Dritten Gesetz zu seiner Änderung (BGBl. 2026 I Nr. 43), neunzehn Befehle an sechs Normen ohne Rest — und ein Gesetzentwurf — das AGG nebst BT-Drs. 21/6178, dreiundzwanzig Befehle an vierzehn Normen, einer zur Prüfung markiert. Der zweite Fall war zunächst als ProdHaftG nebst BT-Drs. 21/4297 gedacht und ist verworfen: Die Modernisierung löst das Gesetz ab, statt es zu ändern, und ergibt einen einzigen Befehl an einer einzigen Norm. Ein Beispiel, das nichts zeigt, ist keines. Mitausgeliefert wird nichts, und die Seite ruft von sich aus nichts ab; sie verweist. Die Verweise öffnen ein eigenes Fenster, weil ein im selben Fenster geöffnetes PDF die bereits gewählten Dateien aus dem Formular nähme. Das Stilblatt bekommt .einfuehrung im Zuschnitt der übrigen Kästen und keine einzige neue Farbe, sodass Hell- und Dunkelfassung ohne Zutun stimmen. Dabei stellte sich heraus, dass gesetze-im-internet.de das Norm-XML gar nicht als solches ausgibt: Es gibt allein …/<kurz>/xml.zip, der unmittelbare BJNR….xml-Pfad ist 404. Eine Einführung, die auf die Fundstelle verweist, wäre ohne Weiteres eine Anleitung zum Entpacken von Hand gewesen. Der Ladepfad nimmt das Archiv deshalb nun unmittelbar an (ZipAuspacker, vorgeschaltet in Pipeline.ladeStammgesetz; DateiTyp kennt die Art ZIP). Die Änderungsdokumente bleiben unberührt — Gesetzblätter und Drucksachen kommen nirgends als Archiv. Das Format wird von Hand gelesen, und zwar aus einem Grund, der außerhalb der Browserfassung sinnlos wäre: Von den Archivklassen des JDK trägt dort allein der Inflater, den InflaterErsatz auf jzlib zurückführt; ZipInputStream, ZipFile und CRC32 beruhen auf nativen Bindungen, die Web Image nicht kennt. Auch die Prüfsumme rechnet die Klasse deshalb selbst. Gelesen wird über das Zentralverzeichnis am Dateiende, nicht über die örtlichen Vorspanne, deren Größenangaben bei nachgestellten Beschreibern erst hinter den Daten stehen. Gewählt wird der einzige auf .xml endende Eintrag: Den Gesetzen mit Anlagen legt die Fundstelle deren Bilddateien mit ins Archiv, es ist also nicht einerlei, welchen man nimmt. Die ZIP-Signatur wird vor der PDF-Signatur geprüft, weil sie am ersten Byte verankert ist, während jene ein Vorschaufenster durchsucht — ein Archiv mit der Zeichenfolge „%PDF-“ in seinen gepackten Daten wäre sonst als PDF angesprochen worden. Die Quellenzeile der Synopse nennt weiterhin die angegebene Datei (uwg.zip), nicht den Eintrag; sie soll den Weg zur Fundstelle zurück beschreiben. Den Eintragsnamen trägt die ausgepackte Quelle gleichwohl, damit Protokoll und Ladefehler das Gesetz benennen und nicht seine Verpackung. Geprüft ist beides gegen die Fundstelle selbst, nicht nur gegen den Beispielkorpus: 312 Prüfungen laufen durch, darunter acht neue zum Auspacken. Das unmittelbar bezogene uwg.zip ergibt auf der Befehlszeile mit dem BGBl-Regelungstext dieselbe Ausfertigung wie die entpackte Fassung — die beiden unterscheiden sich allein in der Quellenzeile —, und dieselben 19/0/6 ergeben sich im neu gebauten Wasm-Modul im Browser, ohne eine Beanstandung in der Konsole. Der Block hält bei 320 Pixeln Breite ohne Überlauf. Unberührt bleiben pom.xml, deploy/webpaket.sh und die nginx-Vorlage: Da nichts mitausgeliefert wird, braucht es weder einen zweiten Ressourcen-Block noch neue MIME-Typen — deren types-Block ersetzt die Zuordnung vollständig und wäre sonst die Falle gewesen. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: I5f6566ca89aa5198452b0b29ac5b38da9eb26137
* Die Umnummerierungs-Kaskade war keine Grenze, sondern drei MängelMatthias Andreas Benkard9 days3-71/+179
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Die mehrschrittige Umnummerierungssequenz galt bislang als wesenseigene Grenze: Ihre Folgeänderungen gingen in den Abschnitt „Manuell prüfen“, im bayerischen Belegfall drei von 154 Anweisungen (48. a) ee), 48. a) gg), 48. b) cc), sämtlich aus der Neunummerierung des Bußgeldkatalogs in Art. 56 BayJG). Die Prüfung gegen die amtliche Nachfassung zeigt, dass es keine Grenze war, sondern drei voneinander unabhängige Mängel. Erst die Diagnose vor jeder Änderung — die Begründungen der drei Fälle ausgeben — hat sie getrennt; die Vermutung, alles sei eine Frage der Reihenfolge, traf nur auf einen von ihnen zu. Geordnet werden nun Schritte statt Befehle (BefehlAnwender). Ein Verbund aus Umnummerierung und Begleitänderung trägt zwei gegenläufige zeitliche Ansprüche: Die Umnummerierung gehört vor denjenigen, der ihre Bezeichnung neu besetzt, die Begleitänderung an ihre Stelle im Dokument, denn sie setzt die vorangegangenen Punkte als vollzogen voraus. Bisher setzte sich der erste Anspruch für die ganze Einheit durch. Bei gg) — „Die bisherige Nr. 11 wird Nr. 12 und die Angabe „schriftliche“ wird gestrichen“ — stand die Angabe vorgezogen noch zweimal im Artikel und war zu Recht mehrdeutig; an ihrer Dokumentstelle hat ee) die zweite Fundstelle längst ersetzt. Die Befehlsliste wird dazu vor der Anwendung zu Schritten aufgefaltet (je Teilbefehl einer); protokolliert wird unverändert je Befehl, ein Verbund gilt nur als angewandt, wenn jeder Teil gegriffen hat. Die Ordnungsregel selbst bleibt wörtlich dieselbe („wer eine Bezeichnung räumt, kommt vor dem, der sie neu besetzt“), wird aber erst jetzt vollständig durchgesetzt. Bisher rückte ein Befehl nur einmal vor seinen ersten Kollisionspartner; auf Schritten zerreißt das die Ketten — von „Nr. 15 wird Nr. 16“, „Nrn. 13 und 14 werden Nrn. 14 und 15“, „nach Nr. 12 wird Nr. 13 eingefügt“ zog die Einfügung nur das letzte Glied vor sich her. An die Stelle tritt eine Tiefensuche auf der Dokumentordnung: vor jedem Schritt erst rekursiv seine Räumer, eine ringförmige Abhängigkeit bricht ab. Kein Kahn mit „frühester bereiter Knoten“ — der zieht unbeteiligte Schritte vor und bricht genau die Begleitänderung wieder. Die beiden übrigen Mängel lagen anderswo: * Das Muster der lokativen Klausel (BefehlErkenner) schloss mit einer Wortgrenze. Hinter einem abgekürzten Bezeichnungswort steht aber schon der Abkürzungspunkt, und zwischen ihm und dem Leerzeichen liegt keine Wortgrenze; der Klausel entging deshalb jede Kurzform („in Abs. 2 …“, „in Buchst. b …“). Die unabgekürzten Formen trafen zu, weshalb es nie auffiel. * Der StellenAufloeser verlangte hinter der Aufzählungsmarke Text auf derselben Zeile. Eine Einheit, die sich vollständig in ihre Untergliederung ergießt, führt ihre Marke allein (Art. 56 Abs. 2 Nr. 12 BayJG, darunter nur die Buchstaben a und b) und galt als nicht auffindbar. Der bayerische Belegfall steht damit auf 154 angewandten Anweisungen ohne Rest (zuvor 151), die Ausschussfassung der Beschlussempfehlung zum GEG auf 68 statt 67 — dort greift nun eine Bereichs-Umnummerierung in § 108, die vorher nur ihr erstes Glied vorziehen konnte; § 108 liegt jedoch in der bekannten Zone, in der das Beispiel-XML eine andere Fassung ist als die vorausgesetzte, taugt also nicht als Beleg. Belegt ist die Ordnung an Art. 29a und Art. 56 BayJG und an den Kaskaden des WDR-Gesetzes. Alle übrigen Bezugszahlen des Bundes und der Länder sind unverändert; 304 Prüfungen laufen durch. Der Akzeptanztest hält zwei Mängel fest, die er nicht behebt. Beide bestanden schon zuvor, was ein Lauf gegen den Ausgangsstand belegt: * Der Block aus „Nach Nr. 4 werden die folgenden Nrn. 5 bis 7 eingefügt“ tritt an das Ende des Absatzes statt hinter die Nr. 4, und der leere Platzhalter „7. (aufgehoben)“ der Altfassung bleibt zwischen den Nrn. 9 und 10 stehen. Dass jede Anweisung greift, heißt eben nicht, dass die Norm in allem der amtlichen Nachfassung gleicht; der Test sagt das ausdrücklich. * Die Heilung doppelter Leerzeichen nach einer Streichung wirkt auf den gesamten Zieltext und verkürzt dabei die Einrückung der Aufzählungszeilen. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: Id61babe60226d27b5c05ff971927d61713b01781
* Die Beschlussempfehlung liefert die beschlossene FassungMatthias Andreas Benkard10 days5-97/+548
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Bisher wurde eine Beschlussempfehlung erkannt und mit Begründung übergangen: Ihre maßgebliche Fassung steht in der zweispaltigen Zusammenstellung, deren rechte Spalte für sich unlesbar ist. Sie druckt Unverändertes nicht ab, sondern vermerkt bloß „unverändert“ — und zwar nicht nur je Gliederungspunkt, sondern auch zeilenweise innerhalb zitierter Blöcke, weshalb ihre Anführungszeichen nicht aufgehen. Eine Auflösung über die Gliederungspfade scheitert daran nachweislich (45 statt 117 Befehlen). Maßgeblich ist die Grundlinie: Beide Spalten sind zeilensynchron gesetzt, jeder Vermerk steht auf der Höhe der Entwurfszeile, die er meint. Der neue `ZusammenstellungsLeser` führt sie über Seite und Grundlinie in eine gemeinsame Lesereihenfolge zusammen und entscheidet dann Zeile für Zeile: „unverändert“ holt den Wortlaut aus der Entwurfsspalte, „entfällt“ streicht ihn, sonst gilt die Ausschussspalte. Das Vokabular der rechten Spalte ist ausgezählt genau zweiteilig — 166-mal „unverändert“, 17-mal „entfällt“. Dabei meint „unverändert“ den Wortlaut, nicht die Zählung: Streicht der Ausschuss einen Punkt, rücken die folgenden auf, und sein „5. unverändert“ steht der Entwurfsnummer 6 gegenüber. Die Marke kommt deshalb von rechts, der Wortlaut von links. Voraussetzung dafür war, die Geometrie aus dem `FontgroessenFilter` herauszuführen. Sonst verließ ihn nur das End-X der Zeile; jetzt trägt jede Zeile Seite, Grundlinie und End-X (`FontgroessenFilter.Zeile`). Die Metadaten reisen weiterhin im Textstrom mit, statt nebenher gesammelt zu werden: Nur so ist ihre Zuordnung zu den Zeilen gesichert, denn PDFBox schreibt Zeilentrenner an mehreren Stellen. Der Schritt ist reine Umstrukturierung, sämtliche gepinnten Zahlen bleiben unverändert. Drei Fehler standen im Weg, alle in der Spaltentrennung: * Der Steg wurde einschließlich der Leerzeichen-Glyphen gemessen, die im Textstrom fast bis an die nächste Spalte reichen. Ein 7-pt-Steg schrumpfte so auf 4,7 pt und galt als durchlaufender Text, was in BT-Drs. 20/7619 die gesperrte Marke des Punktes 14 mitten entzweiriss. Gemessen wird jetzt an den sichtbaren Zeichen; die Schwelle sinkt von 12 auf 6 pt und liegt damit belegt zwischen dem schmalsten echten Steg (6,9 pt) und dem breitesten Wortzwischenraum einer ganzseitenbreiten Zeile (5,4 pt) — in beiden Belegdokumenten übereinstimmend. * Gedrehter Text hat Koordinaten in seinem eigenen Bezugssystem. Der Randvermerk „Vorabfassung – wird durch die lektorierte Fassung ersetzt“ lief mit seiner vermeintlichen Grundlinie quer durch beide Spalten und zerschnitt deren Zeilenfolge. Beim Spaltenauszug bleibt er nun draußen; beim ungeteilten entfernt ihn wie bisher der TextBereiniger. * Die Zeilenenden müssen je Spalte klassifiziert werden, bevor die Spalten zusammenkommen: Ihre Satzspiegelränder liegen verschieden, in der gemischten Fassung fände keine ihren eigenen Ausrichtungs-Cluster wieder und die Silbentrennung bliebe ungeheilt („An- gabe“). Die Quellenzeile weist die verwendete Spalte als „[Beschlussempfehlung …, Ausschussfassung]“ aus. Lässt sich die Zusammenstellung nicht auflösen, wird die Datei nach wie vor mit Begründung übergangen — eine halb aufgelöste Fassung auszugeben wäre schlimmer als keine. Zwei Belegfälle: BT-Drs. 20/7619 (GEG) mit 67 angewandten Befehlen, wobei die Ausschussfassung mehr Befehle trägt als der Regierungsentwurf daneben — der Ausschuss hat zwei Artikel ergänzt —, und BT-Drs. 19/24334 (Drittes Bevölkerungsschutzgesetz) mit 47. Der große manuelle Rest hat denselben Grund wie beim Entwurf: Das Beispiel-XML ist die Urfassung des GEG von 2020, geändert wird eine Fassung von 2023. Die verbliebenen zwei unerkannten Befehle gehen auf Setzfehler der amtlichen Drucksache zurück (fehlender Schlusspunkt, überzähliges Anführungszeichen) und sind zeichengenau so übernommen. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: Id5db19cb8f27a51af0c23cfd4a9969dd4120f569
* Die Web-App rechnet im Browser statt auf dem ServerMatthias Andreas Benkard10 days10-494/+155
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Bisher lief die Weboberfläche auf einem JDK-HttpServer: Uploads landeten als temporäre Dateien auf dem Server, die Synopse entstand dort. Das kostete Betrieb (systemd, Reverse Proxy, Rate-Limiting, Timeouts) und verlangte ein Datenschutzversprechen, das nur zusicherbar, nicht nachprüfbar war — gerade Entwurfstexte verließen den Rechner der Nutzer:innen. Neu übersetzt `./mvnw -Pwasm package` dieselbe Pipeline mit GraalVM Web Image (`native-image --tool:svm-wasm`) nach WebAssembly, PDFBox eingeschlossen. Ausgeliefert werden nur noch statische Dateien; gerechnet wird im Browser. Die erzeugte Synopse ist byteweise identisch mit der der Befehlszeile (SHA-256 verglichen für IfSG 48/27/21 und BayJG 151/3/54). Die Befehlszeile bleibt unberührt: `./mvnw package` erzeugt unverändert das JAR, alle Optionen und Meldungstexte sind gleich, das Wasm-Profil ist rein additiv und verlangt Oracle GraalVM 25.1+ (die CE hat kein Web Image). Portabilitätsschnitt (nützt beiden Fassungen): * `Quelle` (Name + Bytes) ersetzt `Path` in der Pipeline; nur die Befehlszeile kennt noch ein Dateisystem. Der Name trägt genau den bisherigen `getFileName()`-Text, damit Warnungen und Quellenzeile wortgleich bleiben. * `DateiTyp` erkennt PDF/XML/Klartext an den Signaturbytes. Tika entfällt — eine schwergewichtige Abhängigkeit samt ServiceLoader- und XML-Konfiguration weniger, was der Wasm-Übersetzung unmittelbar zugutekommt. Vier Eigenheiten von Web Image, die der Quelltext jeweils an Ort und Stelle vermerkt: * `java.util.zip.Inflater` ist nicht angebunden (GR-65205), ohne Inflate ist kein PDF lesbar. `InflaterErsatz` substituiert ihn durch jzlib. * Typisierte JS-Felder lassen sich nicht nach `byte[]` umsetzen („byteArrayHub is not defined“); der Dateiinhalt wandert als Base64. * JULs Standardformatter ruft `StackWalker`, den es dort nicht gibt. * Im Worker fehlt `document.currentScript`, worauf die Laufzeit das Wasm-Modul neben `worker.js` sucht; die VM wird deshalb mit ausdrücklichem Pfad ein zweites Mal gestartet. Die Reachability-Metadaten stammen aus einem Lauf des Tracing-Agents über die Pipeline; die PDFBox- und FontBox-Ressourcen sind als Globs ergänzt, sonst scheitern PDFs an „Could not find referenced cmap stream Identity-H“. Entfallen: WebMain, UploadHandler, StaticHandler, Multipart und die systemd-Unit. Die nginx-Vorlage liefert jetzt statische Dateien aus, und die Datenschutzseite sagt, was nun stimmt: Die Dateien verlassen den Rechner nicht. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: I38faf2ac0f764d601f080d4276babe4747773683
* QuelltextformatierungMatthias Andreas Benkard10 days16-182/+228
| | | | Change-Id: I51c63959e75af9373625a52b74d1d31f727a67ed
* Quellformate jenseits des verkündeten GesetzesMatthias Andreas Benkard11 days14-155/+1646
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Bisher nahm ÄndGgner nur verkündete Artikelgesetze und Gesetzentwürfe an, ohne den Unterschied zu kennen. Dokumente, die einen Entwurf ändern statt ein Gesetz, lieferten still null Befehle. Neu: * Dokumentart-Erkennung aus dem Rohtext (DokumentErkenner, DokumentArt, DokumentKopf) — vor der Bereinigung, weil gerade die Drucksachenköpfe als Kolumnentitel wegfallen. Erkannt wird nie am Dateinamen: Die Beispieldaten enthalten eine Datei namens Beschlussempfehlung, die in Wahrheit ein Entschließungsantrag ist. Dokumente ohne Befehle erzeugen jetzt eine benannte Warnung statt stiller Leere. * Änderungsanträge (AenderungsantragParser, EntwurfsPatcher): Der Antrag wird auf den Entwurfstext angewandt, der geänderte Entwurf danach wie gewohnt auf das Stammgesetz — eine komponierte Synopse. Drucksachen- und Gesetzesstellen bleiben getrennte Typen; die Zuordnung Antrag→Entwurf läuft über die Drucksachennummer, nicht über die Argumentreihenfolge. Der 1910 Zeilen große BefehlAnwender bleibt unberührt. * Beschlussempfehlungen werden erkannt und mit Begründung übergangen, die den brauchbaren Entwurf beim Namen nennt. Die Spaltentrennung der Zusammenstellung ist gebaut (FontgroessenFilter.Spalte, koordinatenbasiert am Steg) und belegt: Die linke Spalte ergibt Befehl für Befehl den Regierungsentwurf. Die rechte Spalte bleibt bewusst offen — sie vermerkt „unverändert“ auch zeilenweise innerhalb zitierter Blöcke, ihre Anführungszeichen gehen daher nicht auf; nötig ist die zeilenweise Zuordnung beider Spalten über die gemeinsame Grundlinie. Die Synopse kennzeichnet eine Entwurfsfassung als solche. 296 Tests grün (vorher 269); alle gepinnten Akzeptanzzahlen unverändert. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: I0001d25fcbd9969fc06eea675edc4326ce2e02f9
* Add a public web front end alongside the CLIMatthias Andreas Benkard2026-08-096-77/+539
| | | | | | | | | | Extracts the CLI pipeline into a reusable Pipeline class and exposes it via a dependency-free JDK HttpServer (upload form, bounded worker pool with 503 on overload, per-request size/time limits, no persisted uploads), plus nginx/systemd deploy templates and Impressum/Datenschutz placeholders for public operation. Change-Id: I6e8e7afa3c4b1082cdf9e82da0fae0b5b49470ad
* Add Hesse as an extraction and recognition caseMatthias Andreas Benkard2026-08-022-9/+79
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The Hesse GVBl issue 5/2026 amends the ordinance on traffic-law competences. All 21 commands of its article 1 are now recognised. The blocked letterspacing that was expected to be the hard part turns out to affect only the signature block, not a single command. What did need work was three other things: the running GVBl footer and the masthead remainder are now stripped as page furniture, and so is the heading's asterisk footnote ("* Ändert FFN 61-60"), which sits at the foot of a page between two enumeration items and stuck to the command above it. Two command forms are new, both general rather than Hessian: "am Ende" may be absent from a punctuation replacement — the definite article then requires the unit to carry exactly one such mark, which the applier checks — and the object after "durch" may be elided in the form without a location, as it already could in the form with one. Where a punctuation replacement follows a renumbering in the same sentence it binds to the renumbered unit: unlike a word operation it has no distinguishing target text, so resolving it norm-wide would be useless. A full acceptance test stays out of reach: hessenrecht.hessen.de is the same login-only juris application as the Schleswig-Holstein and Berlin portals. Checking landesrecht-bw.de too, that now covers all four portals that could have supplied an annex-bearing statute, so the LandesRechtTextParser's missing annex headings will need a stem from recht.nrw.de instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: Ief04164ee25653345a011f3ffbbfaa6df1f5c64b
* Run the NRW mass test: WDR-Gesetz and LMG NRWMatthias Andreas Benkard2026-07-286-71/+687
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The 22nd Rundfunkänderungsgesetz amends five laws in one gazette issue. Article 3 (TMZ-Gesetz) was already an acceptance case; this adds the two big ones — Article 1 (WDR-Gesetz, 101 commands across 31 norms) and Article 2 (LMG NRW, 18 commands). Both stem versions come from recht.nrw.de, whose URLs address a version by its date of entry into force, so the results could be checked norm by norm against the official 1 April 2026 consolidations. Applying them turned up work at every layer: Order of application. Renumbering always refers to the original count, not to the state after the preceding items, so a cascade („Absatz 3 wird Absatz 4“, „Absatz 4 wird Absatz 5“, …) applied in document order leaves the law carrying two units of the same name. BefehlAnwender now derives, for each command, which designations it vacates and which it newly occupies, and pulls a vacating command ahead of the first one it collides with. That single rule subsumes the previous §-only special case, orders a cascade descending, and — inside a Sammelbefehl too — keeps a range renumbering from colliding with itself. Bayern's Art. 29a sequence now resolves as well (5 manual residues down to 3). Renumbering an enumeration item. Unlike an Absatz, a Nummer carries its designation as a marker in the text; the applier silently reported success without touching it. It now rewrites the marker and lets a struck placeholder with the target designation give way. Sentence splitting. Gazette citations („(BGBl. I S. 1982)“) and line-initial enumeration markers were splitting sentences, so Satz locators pointed into the wrong place; a sentence may also open with §. Command language: „wie folgt neu gefasst“ and the missing „wird“, a split location („In Absatz 2 wird Satz 1 wie folgt gefasst“), an insertion behind a location prefix, the standalone „Der Punkt am Ende wird durch … ersetzt“, „Dem Wortlaut des Absatzes 3 …“, the verb frame „Es werden ersetzt:“ whose items carry only the location, coordination at „sowie“ and before a punctuation clause, and a stray closing quote at the end of a sentence. Stem format: LandesRechtTextParser reads an „Inhaltsübersicht“ line as the norm of that name (which is what the Angabe commands address) and recognises keyword-less Roman section headings. Everything the two acceptance tests do not apply is named: one command in the WDR-Gesetz whose „Punkt am Ende des Satzes“ is not unambiguous in a two-sentence Nummer. The remaining differences from the official consolidations are places where the official text itself departs from the command wording; SOURCES lists them. Change-Id: I261cf89cd5820f7597bc5db734f322d94b84828e
* Recognise word-anchored insertions; close quotes at the next itemMatthias Andreas Benkard2026-07-286-3/+288
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The documented next step was the word-anchored structural insertion ("Vor den Wörtern „Aus dem Bereich Verkehr:“ wird folgender Absatz 5 eingefügt", Berlin Artikel 1 Nr. 2 b) bb)). StrukturEinfuegung now carries an optional WortAnker, the recogniser has a pattern for the form — placed before STRUKTUR_EINFUEGUNG, which bails out of recognition entirely when the StellenParser cannot read "den Wörtern «0»" — and the applier positions the new unit at the anchor line instead of a structural boundary. Two blockers turned up next to it, both measured on the corpus rather than assumed: The command was not cleanly recognisable at all. Its closing quotation mark is missing in the official text, so the quote swallowed items cc) and c); an anchor alone would have made the command look applicable with a huge foreign quote as its payload. The docs called a boundary at enumeration markers impossible, and that holds for the unguarded form: a quoted amendment provision carries command language itself. Guarded by four conditions together — the article's quotation marks demonstrably do not balance, the next line's marker is at the same or a shallower level than the one the quote opened on, that line carries command language, and closing here leaves the rest of the article balanced — it fires exactly twice in the whole sample corpus, both times where the missing mark belongs: Berlin before cc), and GV. NRW. Artikel 2 before 12. Without the guards it also fires inside quoted amendment provisions in the BayJG and GModG documents; the balance gate rules those out. The GV.-NRW. gazette encodes part of its umlauts decomposed (u + U+0308, 79 places, dozens of them "eingefügt:"/"angefügt:"). Invisible in the text, but a different word to every command pattern, so those commands could never match. TextBereiniger now normalises to NFC first — not NFKC, which would flatten the official sentence numbers ¹²³ and take SatzTeiler and Superskript their basis. Every other sample document, every gii-XML stem and every hand-kept plain-text stem is already NFC, so the pinned figures stay bit-identical. Berlin Artikel 1 is thereby fully recognised (6 of 6 commands, none unknown). One pinned statement changed for a stated reason: the NRW acceptance test asserted the article-heading warning, which is now preempted by the item boundary inside Artikel 2 — Nr. 12 survives instead of being swallowed. 241 tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: Ic579bec6fecc2495a61b3dcf87c9ad71b42f34ae
* Support unnumbered named Anlagen; strip control charactersMatthias Andreas Benkard2026-07-262-1/+31
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Investigating "Anlagen as amendment target" corrected the diagnosis from yesterday: Anlagen are supported. In gii-XML they are norms of their own ("Anlage 8", "Anhang"), Stelle.anlagenEnbez() resolves them, and the GEG sample applies annex-targeted commands — only two of its fifty manual entries touch annexes, both for content reasons. What actually failed in Berlin was narrower, and is fixed: * A law with a single annex names it after the provision it belongs to ("Die Anlage zu § 2 Absatz 4 Satz 1"). StellenParser demanded a number after "Anlage" and rejected the phrase, so the article's frame command went unrecognised and its items lost their context. The annex now carries the designation "Anlage", like "Anhang". The naming suffix is skipped deliberately: read along, the "§ 2" inside it would become the target and the wrong norm would be amended. * C0 control characters are now stripped. Berlin's enumeration indent carries a U+0007 that sits invisibly in front of the command and made its start-of-line anchor miss. * Command verbs torn apart by a stray space in the official typesetting ("ein gefügt" for "eingefügt") are rejoined, restricted to sequences that are not valid German word order. Three of Article 1's four commands are now recognised. The fourth places the new unit by a word anchor ("Vor den Wörtern … wird folgender Absatz 5 eingefügt"), which StrukturEinfuegung cannot express; that and the inability of the plain-text stem format to carry an Anlage at all are recorded in Landesrecht-Beispiele.adoc. 231 tests green; all pinned figures unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: Ic5804bf343dd5d8ce435c07b8f882ca0372890d2
* Berlin column spike: stream order already reads correctlyMatthias Andreas Benkard2026-07-262-4/+17
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The plan assumed the two-column Berlin gazette would need x-coordinate column detection. It does not: the content stream already emits the left column fully before the right one, so commands come out in reading order. The test pins that finding. What does go wrong there is different, and two parts of it are fixed: * The article's amendment formula carries its own target ("§ 2 Satz 1 des Gesetzes zur Errichtung … wird wie folgt geändert"), but the following items never inherited it. The pattern's leading \b could never match before "§" — a non-word character needs a word character in front of it — so the paragraph branch was dead code and only the Bavarian "Art." branch ever fired. * The running page header is not a line of its own; it sits inside the body text of the following column, so it has to be cut out rather than dropped as a column title. Two blockers remain and are recorded rather than papered over: Article 1 amends an *Anlage* with its own numbering, which the model does not carry (the same gap that shelved Baden-Württemberg), and full-width frames (title block, imprint) sit elsewhere in the stream than on the page, which would need a real XY-cut. Stems are unobtainable either way — gesetze.berlin.de is an authenticated juris app like the Schleswig- Holstein portal. 229 tests green; all pinned figures unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: I4619dbe96a023d61104d7107a9a34533123d1bc5
* Support the two Lower Saxony insertion forms; NEFG now applies fullyMatthias Andreas Benkard2026-07-253-7/+75
| | | | | | | | | | | | | | | | | | | | | | | | | | | | The NEFG acceptance case left two commands as "check manually". Both are now supported, so Lower Saxony applies 7 of 7 with no remainder. * Lower Saxony cites inserted paragraphs with a space before the sub number ("§ 2 a"); the rest of German legal usage writes "§ 2a", which is what the Stelle and level parsers expect. TextBereiniger now pulls the form together, leaving enumeration markers of the amendment act alone ("… nach § 8 c) In Absatz 2 …"). * "In Kapitel 4 wird nach § 12 der folgende neue § 13 angefügt" — the division only names the section the new paragraph lands in; the anchor governs the position, as in the plain form. The optional "neue" is now accepted in both. Applying that insertion needed one more thing: the paragraph it creates is only vacated by the *following* command ("Der bisherige § 13 wird § 14"), so in document order the two collided and the insertion was rejected. "Bisherig" denotes the state before the amendment, so the renumbering logically precedes the reoccupation and is now pulled ahead. The log still lists commands in the order of the amendment act. 226 tests green; all other pinned figures unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: I0e2ea5b0389f8002157eeb8471800da51febe739
* Add North Rhine-Westphalia TMZ-Gesetz acceptance caseMatthias Andreas Benkard2026-07-252-3/+48
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Article 3 of the 22. Rundfunkänderungsgesetz (GV. NRW. 2026 S. 202) amends the Telemedienzuständigkeitsgesetz. All four commands apply cleanly, and the result matches the official consolidated version of 1 April 2026. recht.nrw.de turned out to be the one remaining state portal that serves consolidated laws as server-rendered HTML, with the version history encoded in the URL date prefix — so the pre-amendment stem (27 Apr 2022) was directly obtainable. Two fixes were needed: * LandesRechtTextParser read the parenthesised title suffix wholly as the abbreviation. State laws routinely carry both short title and abbreviation there ("(Telemedienzuständigkeitsgesetz – TMZ-Gesetz)"), and the amendment act cites the short title, so article selection never matched. The parser now splits at the dash. * ZitatExtraktor swallowed the rest of the document when a closing quotation mark is missing in the official text (Article 2 no. 11 here), which hid Articles 3 to 5 entirely. An amendment act never quotes across an article heading, so an open quotation is now closed before a bare "Artikel N" line, with a warning. Quoted paragraphs are unaffected: Bavarian norm heads read "Art. N". 223 tests green; all previously pinned figures unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: I78470075d9be2433966b821b6fe1e21808996c46
* Add Schleswig-Holstein extraction case and three cleanup fixesMatthias Andreas Benkard2026-07-251-3/+40
| | | | | | | | | | | | | | | | | | | | | | | | | | | | The GVOBl. Schl.-H. amendment act (Kommunalrecht, 2026/27) brings three conventions the tool had not seen: * a running two-line page header that sits between articles and glued itself onto the following article heading, * superscript footnote markers on article headings ("Artikel 1 ¹)"), which defeated the article split, and * ragged-right typesetting, where the geometric margin detection finds no alignment cluster and reports full lines as hard line ends, so a trailing hyphen was never joined. All three are fixed in TextBereiniger; the hyphen rule now trusts the extractor's trailing-space signal over the geometric class. A full acceptance test is not possible: the Schleswig-Holstein law portal serves document bodies only through an authenticated API, and the last freely archived consolidated Gemeindeordnung (14 Nov 2022) predates the amended § 34a. The new test therefore covers everything up to and including command recognition; the blocker is recorded in Landesrecht-Beispiele.adoc. 219 tests green; Bund/Bayern/Sachsen/Niedersachsen figures unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Change-Id: I9b3ab8a756372504b51c4d410b491fded255da96
* Add Lower Saxony NEFG acceptance case and two parser robustness fixesMatthias Andreas Benkard2026-07-242-2/+15
| | | | | | | | | | | | | | | | | | | | | | | | | | | Second §-structured Landesrecht case for wave six: the Lower Saxony ELER Support Act (NEFG, unchanged since 2022, derived from the two-column Nds. GVBl 2022 No. 33) plus the amending act of 28 Jan 2026 (Nds. GVBl 2026 No. 10). Five of seven commands apply — both "erhält folgende Fassung" re-enactments (§ 1(1), § 2), the citation replacement in § 1(5), the words insertion in § 6(1), and the § 13 → § 14 renumbering. - LandesRechtTextParser: a norm head no longer accepts a title ending in a full stop, so a cross-reference sentence at line start ("§ 7 GAPInVeKoSG findet entsprechend Anwendung.") is not misread as a norm head "§ 7" (which had made the monotonicity check swallow the norms in between). - TextBereiniger: strip the Lower Saxony GVBl running footer and publisher line. The footer sat between two enumerated items and otherwise attached to the preceding re-enactment quote, defeating its end-of-sentence anchor. - Add sampledata/Niedersachsen/NEFG-alt.txt and EndToEndTest.nefgAcceptance. Two insertion forms remain documented residues (manuell prüfen): inserting a §-block with a space-separated sub-number ("§ 2 a") and a chapter-scoped §-block insertion ("In Kapitel 4 wird nach § 12 … angefügt"). 221 tests green; federal, Bavarian and Saxon pinned numbers unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Change-Id: Ie23eee1f596ec8eecc6a73f6ce585a22e73ec6c9
* Add Saxony SächsBeamtVG acceptance case (first non-Bavarian Landesrecht)Matthias Andreas Benkard2026-07-242-1/+28
| | | | | | | | | | | | | | | | | | | | | | Complete the first §-structured Landesrecht acceptance case for wave six: the Saxon Civil Servants' Pension Act (SächsBeamtVG) in its 1 Feb 2025 version plus the amending act of 13 April 2026 (SächsGVBl. S. 134). All five commands are recognised and applied without residue — the amount replacements in § 47 and the re-enactment of § 80(1) sentence 2. - LandesRechtTextParser: recognise arabic, keyword-first structural headings ("Abschnitt 1", "Unterabschnitt 2", "Teil 3") via GLIEDERUNG_ARABISCH, in addition to Bavaria's roman "I. Abschnitt" form. Bavaria behaviour unchanged. - TextBereiniger: strip the footer of revosax full-text output (REVOSAX_FUSS, "https://…revosax… Fassung vom … Seite N von M"). - Add the consolidated pre-amendment stem as canonical plaintext (sampledata/Sachsen/SaechsBeamtVG-alt.txt, derived from the archived revosax full text) and EndToEndTest.saechsBeamtVgAcceptance (5 applied, 0 manual). 220 tests green; federal and Bavarian pinned numbers unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Change-Id: I8e878bc9506f24028cf643a873de000f3d066d64
* Generalize Bavarian loader to a shared Landesrecht loaderMatthias Andreas Benkard2026-07-195-51/+102
| | | | | | | | | | | | | | | | | | | | | | | | | | | Prepare wave six (Landesrecht of the remaining Länder) by lifting the Bavaria-specific stem-law loader to a shared one. Unlike Bavaria (Art./§ inversion), the other Länder structure their stem laws in §, so the loader must handle both sigils. - Move gesetz/bayern/{BayRechtLoader,BayRechtTextParser} to gesetz/land/{LandesRechtLoader,LandesRechtTextParser}; the norm head now matches "§ N" as well as "Art. N" and derives the sigil per norm from the match. Cross-reference keywords in the norm head are excluded only as whole words (so a title "Satzungen" no longer trips on "Satz"), and the juris abbreviation may be multi-token ("(GO NRW)"). - Derive the superscript mode data-drivenly from the loaded stem law (Superskript.traegtSatznummern) instead of from "is it gii-XML": Bavaria and Lower Saxony keep their amtliche Satznummern, the Bund and Länder without official sentence numbering drop them. - Recognize the neufassung idiom "erhält/erhalten folgende Fassung" (Schleswig- Holstein, Niedersachsen) via a NEUFASSUNG_VERB building block, additive to "wird/werden wie folgt gefasst". Federal and Bavarian behaviour is unchanged (219 tests green, pinned acceptance numbers UWG 19/0, GEG 66/53, GEG-GModG 90/9, IfSG 42/24, BayJG 149/154 hold). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Change-Id: I55a70a7932bf657a2346ca70f3fa05e173bf80ad
* Support Bavarian state law (BayJG) with Art./§ inversion and superscriptsMatthias Andreas Benkard2026-07-1918-105/+1412
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Extend ÄndGgner from federal-only to Bavarian state law. Bavarian base laws are structured in Artikel while their amending acts are structured in Paragraphen (inverted from the Bund); official sentence numbers and footnote markers are carried as Unicode superscripts (¹²³, ⁶)). - Generalise Stelle.Paragraph to carry a sigil ("§" or "Art.") and route it through recogniser, applier and resolver (bit-identical Bund output). - Superscript pipeline: geometric detection in FontgroessenFilter (SuperskriptModus BEHALTEN), Superskript util, exact sentence splitting in SatzTeiler, label-based sentence resolution. - BayRechtLoader/BayRechtTextParser for gesetze-bayern.de PDF/plaintext. - §-structured amending acts, GVBl/Landtag column titles, non-breaking spaces, the GVBl continuation quote; Bavarian command forms (footnote aufhebung, Satznummerierung streichung, Wortlaut forms, Halbsatz, gapping chains). Acceptance (EndToEndTest.bayJgGvblAcceptance): the pre-2026 BayJG fassung (BayJG-alt.txt, reconstructed from Wayback single-article snapshots) with GVBl 6/2026 §§ 1-2 applied — 154 commands, 0 unknown, 149 applied automatically, 5 pinned residuals (follow-up edits inside two multi-step renumbering sequences in Art. 29a and Art. 56). Application-side fixes surfaced by the acceptance run (all Bund-safe): sentence-start superscript before §; sentence boundaries = {0} ∪ {each number ≥ 2}; gapping scope inheritance for bare word operations; footnote definition lines hidden from word operations; absatz aufhebung marks "(weggefallen)" keeping its number, and absatz renumbering overwrites an empty placeholder (weggefallen/gegenstandslos) target. 211 tests green (mvnw verify); federal reference numbers unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Change-Id: Iebae4c17ca90755c5fd36251362042f3d5796fd0
* Classify line ends by alignment clusters and canonicalize definition listsMatthias Andreas Benkard2026-07-193-54/+97
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The percentile-based right-margin estimate from the previous commit misclassified two real-world layouts: at the column switch of the two-column old-BGBl format a window contains both column margins, so the 90th percentile returned the right column's margin and marked every left-column soft wrap as deliberate (vetoing hyphen mending, IfSG 2020: "Impf- surveillance"); and series of equally wide centered "Artikel N" headings in draft bills formed the top of their window's distribution, so they counted as full-width and were reflowed into the following line, breaking teileInArtikel's line anchor (ProdHaftG RegE). Classify against alignment clusters instead: a line end is soft when at least five window lines end within 4 pt of it (the justified block's own margin), hard when such a cluster runs at least 10 pt above it, unclassified otherwise. Structure anchors ("Artikel 2", "§ 19", bare Gliederung labels) are always exempt from reflow since equal-width heading series still form sham clusters. Also flush the pending line end-X at page ends — without it, the first line of the next page inherited the previous page's footer geometry and was misclassified (UWG: "Artikel 2" swallowed the closing provisions). Markerless joins additionally never cross into a following enumeration marker line ("...vorgesehen und" + "d) die Überwachung" no longer glues to "undd)"), a text-level veto that also works without geometry. On the base-law side, ContentFlattener now separates sibling <LA> elements within a <DD> — the short-label/definition pairs of the UWG Anhang and § 2 IfSG were previously glued without any separator ("Irreführung über Unternehmereigenschaftdie unwahre Angabe...") in both synopsis columns. The continuation line is indented two spaces deeper than its enumeration line so StellenAufloeser.zeilenBlock keeps it inside the unit's block ("Anhang Nummer 31 Buchstabe b"). BefehlAnwender writes the same canonical shape when inserting or recasting quoted units (rueckeZitatEin), so XML-derived and PDF-derived items agree: " 2a. Stichwort" + " Definitionstext". Verified by mvnw verify (179 tests, 7 new) and a before/after sweep of all 21 sample law/amendment combinations: applied/manual counts are unchanged throughout (UWG 19/0, GEG-BGBl 66/53, GModG-RegE 77/20, IfSG-0645 42/24, ...), no marker characters leak into the HTML output. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Change-Id: Ib7faae1bc5c59bda83f648a579af16a4c77e4285
* Classify line breaks geometrically via PDFBox coordinates and support list ↵Matthias Andreas Benkard2026-07-174-48/+317
| | | | | | | | | | | | | | | | | | | | | | | indentation Replaces the heuristic character-count full-width check with real geometric layout analysis. FontgroessenFilter collects the right-edge X coordinate of the text runs on each page and classifies line endings against the local 90th percentile of page margins (using a 20-line sliding window) into soft, hard, or unclassified breaks. TextBereiniger uses this classification to reflow only soft line wraps (WEICH) and keep deliberate ones (HART or UNBEKANNT). Also aligns quote normalization and GII-XML flattening to indent continuation lines in lists (e.g., hanging definitions in UWG Anhang): - BefehlAnwender.normalisiereZitatText applies to single-unit Neufassung and indents continuation lines deeper (4 spaces) than list items (2 spaces). - ContentFlattener generates the same structure for sibling <LA> elements inside <DD>. This ensures the paragraph parser correctly groups these lines as child lines. Change-Id: I1fd7039d4933a273cc2dc55f958d34b439e2302c
* Fix markerless line-break join gluing whole words without a spaceMatthias Andreas Benkard2026-07-172-3/+53
| | | | | | | | | | | | | | | | | | | | | | | | TextBereiniger.verbindeUmbrueche assumed any letter-ending, no-trailing-space line followed by a lowercase continuation was a hyphen-less mid-word split (PDF extraction artifact) and joined it with zero separator. That assumption breaks for deliberate word-boundary breaks, e.g. the short-label/hanging- indent definition format in the UWG Anhang ("...Nachhaltigkeitssiegels" + "das Anbringen..."), producing glued words in the rendered synopsis. Gate the markerless join on the candidate line reaching a locally-typical "full column width" (90th percentile in a ±20-line window), since automatic wraps always land near the column edge while deliberate breaks don't. The window is local rather than document-wide because some source PDFs mix column widths within one document (narrower Regelungstext vs. wider Begründung), which a global statistic would otherwise penalize. Also route the single-unit Neufassung fallback through the existing normalisiereZitatText normalization, matching its sibling code paths, so that internal line breaks now more often preserved by the fix above don't leak into the HTML output as spurious line breaks instead. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Change-Id: I0a79f22f9286cd4a90eea32bad9b54cb4a082cf4
* Support Anhang/Anlage targets, Inhaltsübersicht Angabe commands, and ↵Matthias Andreas Benkard2026-07-1617-196/+1859
| | | | | | | | | | | | | | Gliederungs-Überschriften insertion/replacement Fourth enablement wave: Anhänge/Anlagen resolve as ordinary norm targets with nested Nummer/Buchstabe block resolution, InhaltsuebersichtAnwender applies Angabe commands automatically, GliederungsUeberschriften handles inserting and replacing structural headings, the law's own heading can be recast, and FontgroessenFilter now determines body text per page instead of document-wide. Sample data (UWG/AGG/ProdHaftG) now applies with 0 manual cases. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Change-Id: I1a76750afbecfa1aae734bcdabc91470a28420ca
* Support range structure replacement, struct renumbering, and multi-paragraph ↵Matthias Andreas Benkard2026-07-156-29/+404
| | | | | | | | | | | | blocks Adds bisStelle to StrukturErsetzung for coordinated target ranges, recognizes deletion/renumbering of whole structural units (paragraphs and Gliederung entries), and handles insertion/replacement of multi-paragraph blocks split on §-headings. Updates FASSUNGEN.txt accordingly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Change-Id: I637cc5effbe32c094e62c070b63c470149a77c85
* Model and apply structural (Teil/Abschnitt) and TOC commandsMatthias Andreas Benkard2026-07-1310-15/+293
| | | | | | | | | | | | | | | | | | | | | | | | Introduce the Gliederungsbaum into the model and support the structural command families that were the bulk of the remaining unknowns: - Gliederung gains a kennzahl; Gesetz carries the ordered list of Gliederungseinheiten, which the loader now collects. - New Stelle components Gliederungseinheit (Teil/Abschnitt/Unterabschnitt/ Anlage/…) and Absatzbezeichnung; StellenParser parses these plus "Überschrift von <Gliederung>" and drops "Satzteil/Angabe vor Nummer N" chapeau qualifiers. - Gliederungs-Überschrift Neufassung/Streichung apply to the tree and render as a "Geänderte Gliederungs-Überschriften" diff section; Absatzbezeichnung-Streichung removes an Absatz number; Inhaltsübersicht "Angabe(n) zu …" commands are recognized (applied via the existing TOC path). Reuses Neufassung/Aufhebung with structural Stellen rather than adding new command types. Final UnbekannterBefehl counts: GEG 50->5, IfSG 11->3, AGG 2->0 (§ 1 Alters->Lebensalters now applies), UWG 2->0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Change-Id: I0cc91bfe65798140dd982b19f1884dfe60be87c4
* Support range, renumbering, and punctuation commandsMatthias Andreas Benkard2026-07-134-26/+281
| | | | | | | | | | | | | | | | | | | | | | | | | Extend the parser and applier for several command families that were falling back to UnbekannterBefehl: - Range/coordinated Aufhebung ("Die Nummern 1 bis 3 werden aufgehoben.", "Die Absätze 4 und 5 werden aufgehoben.") via bis-range expansion in StellenParser.parseMehrfach and ausStellen. - Range renumbering without "zu den" and for Nummern/Buchstaben ("Die bisherigen Nummern 4 bis 6 werden die Nummern 8 bis 10."); single renumbering now also covers Nummer/Buchstabe and "Die bisherige". - §-range Neufassung ("Die §§ 52 bis 56 werden wie folgt gefasst: …"), splitting the quoted block at "§ N" boundaries. - "Der Wortlaut wird Absatz N." — a new WortlautZuAbsatz command that numbers the previously unnumbered body. - Word-to-punctuation replacement ("… das Wort „oder" am Ende durch ein Komma ersetzt") and comma+words insert/replace variants. StellenParser gains plural component words (Absätze/Sätze/Nummern/ Buchstaben) and bis-range expansion. Cuts UnbekannterBefehl further: GEG 36->21, IfSG 7->3. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Change-Id: I4972d4eb1d6062521a55710c1f23822e8e22b0c4
* Recognize compound and multi-pair amendment commandsMatthias Andreas Benkard2026-07-132-11/+234
| | | | | | | | | | | | | | | | | | | | | | Split BefehlErkenner into a single-command pass and a fallback that composes Sammelbefehle from commands chained with "und"/", wird": - Multi-pair replacement ("… A durch B und C durch D ersetzt") emits one Ersetzung per pair, crossed with the (possibly coordinated) Stelle. - Verbund splitter probes each "und"/", wird" boundary; when both halves parse — trying the right clause as-is, capitalized, or with the left clause's locative prefix — it folds them into one Sammelbefehl. The single pass no longer short-circuits when a pattern matches but its Stelle is unparseable, so the fallback still gets a chance. Add "ein Komma eingefügt" and anchor-first insertion patterns the splitter needs, and let StellenParser.parseMehrfach inherit the component type for bare-number continuations ("Absatz 1 und 5"). Cuts UnbekannterBefehl counts: GEG 50->36, IfSG 11->7, AGG/UWG 2->1. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Change-Id: I8f76edd0c9ed6cf68c10c50f535a8182e13e378e
* Support multi-target "jeweils" commands and range renumberingMatthias Andreas Benkard2026-07-134-14/+196
| | | | | | | | | | | | Add a Sammelbefehl command that applies one operation to several "und"/"sowie"/comma-coordinated Stellen sharing a common prefix, and resolve "Die bisherigen Absätze X bis Y werden zu den Absätzen X' bis Y'" into descending single renumberings. Parsing gains StellenParser.parseMehrfach; the applier folds sub-commands into one log entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Change-Id: Iaa2c7161066d40d0605c00d3b6ca9a08696070e7
* Support the digital BGBl format and bill drafts as inputs.Matthias Andreas Benkard2026-07-139-70/+481
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The Bundesgesetzblatt has been published digitally via recht.bund.de since 2023 in a new single-column layout, and pending amendment acts are only available as Referenten-/Regierungsentwuerfe or Bundestag printed papers. Both now work as patch inputs: * PDF extraction filters out small print by font size (two-pass PDFTextStripper): in the new BGBl format, footnote blocks and superscript footnote markers would otherwise land in the middle of the statutory text, even inside quoted passages. * TextBereiniger recognizes the new page headers (BGBl "Seite N von M", draft page markers " - N - ", Bundestag printed-paper headers) and joins markerless end-of-line hyphenation (BT-Drs PDFs break words without a hyphen character; regular wraps carry a trailing space, so its absence is the signal -- gated on the source using the trailing-space convention at all, protecting hand-written plain-text inputs). BMJV draft templates draw the hanging opening quote after the paragraph marker ("(1) „" / "§ 19„"); this inversion is repaired. * Drafts embed the statutory text between a cover sheet and a Begruendung section; article scanning now stops at the Begruendung heading. Articles without numbered items (single-command articles like ProdHaftG-RegE Artikel 2) are parsed from the preamble rest. * Target-law matching is declension-tolerant ("Das Allgemeine Gleichbehandlungsgesetz" matches "Allgemeines Gleichbehandlungsgesetz") via rough word-stem comparison. * New command form StrukturErsetzung ("§ 2 Absatz 2 wird durch die folgenden Absätze 2 und 3 ersetzt"); "durch die folgende Überschrift/den folgenden § N ersetzt" map to Neufassung; plural insertions ("die folgenden Absätze 6 und 7"), triple-letter outline markers (aaa), the compound punctuation replacement ("durch ein Komma und die Wörter ... ersetzt"), and Inhaltsuebersicht-Angaben inside a context frame are recognized. End-to-end smoke tests cover the four new datasets: UWG (new BGBl format, footnote-filter assertion), GEG 2023 ("Heizungsgesetz", 121 commands), AGG (BT-Drs draft against an unconsolidated base -- the first dataset with real diffs, spot-checked against the official BMJV synopsis), and ProdHaftG (draft whose Artikel 1 is a replacement law and only Artikel 2 amends the base). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Change-Id: Id1b12bc0bee4178bd1a5c55b3a83f6e13944af69
* Implement the synopsis pipeline for amendment acts.Matthias Andreas Benkard2026-07-1324-35/+3187
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Add the full vertical slice from input files to a two-column HTML synopsis: * gesetz, gesetz.gii: immutable law model (Gesetz/Norm/Absatz) and an offline loader for the gii-norm XML format of gesetze-im-internet.de, flattening DL enumerations, tables, and pre blocks. * aenderung: amendment command model as a sealed interface with records (Ersetzung, Neufassung, Einfuegung, Anfuegung, Aufhebung, Streichung, Umnummerierung), including UnbekannterBefehl as the mandatory fallback -- unrecognized commands are reported, never dropped. * aenderung.parse: text extraction (PDFBox in content-stream order for two-column BGBl PDFs, plain text as escape hatch), BGBl header and hyphenation cleanup, quote extraction with placeholder substitution (quoted blocks cannot confuse the command regexes; unbalanced quotes -- which occur in real BGBl documents -- become warnings), outline scanning with successor-validated markers (1., a), aa), aa1)), and the command recognizer with context-frame stacking. * anwendung: sequential command application with a per-command protocol (ANGEWANDT / MANUELL_PRUEFEN plus reason), scope resolution down to sentence/enumeration ranges via a German sentence splitter. * synopse: norm pairing, word-level diff via java-diff-utils, and a self-contained HTML renderer with a Manuell-pruefen section. The CLI gains -o/--output, --vollstaendig, --artikel, --extract-only, and hidden debug flags (--dump-gesetz, --dump-befehle). On the IfSG sample (Drittes Bevoelkerungsschutzgesetz, BGBl. I 2020 S. 2397), 63 of 75 commands in Artikel 1 and 2 parse into typed commands; the remainder are ranges and compound commands that are deliberately out of scope for v1. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Change-Id: I93ef8ba300f1cac70916722f23c0cefc5f62da2b
* Configure Tika correctly.Matthias Andreas Benkard2020-12-062-12/+47
| | | | | | | | | With the new configuration, Tika can now extract text from PDFs and XML documents. Also configures logging for the application. Change-Id: I7a89c2b232ed4e220665dd335a5f5a0cc3ef2994
* Add Apache Tika, remove iText7.Matthias Andreas Benkard2020-11-232-6/+26
| | | | Change-Id: I9736779cf57050a1dfd43d18625eb464a7179a9f
* Modularize, avoid shade plugin by default.Matthias Andreas Benkard2020-11-231-0/+3
| | | | Change-Id: I704d78b7553f057f73e1387078623d8fe5d0b154
* Define CLI arguments.Matthias Andreas Benkard2020-11-221-1/+13
| | | | Change-Id: I37d6cbb000df63013277fe4628cdb6570d68101a
* Project skeletonMatthias Andreas Benkard2020-11-221-0/+20
Change-Id: I5609adfca13ffad643a3db93e848fc01636b066a