LLMs for Clause Rectification in German Employment Contracts
Large Language Models (LLMs) hold the promise of transforming automated contract review from a risk identification task into a remediation process. However, effectively moving from detection to automated rectification presents a distinct challenge, especially in the context of highly regulated General Terms and Conditions: a model must satisfy complex legal validity constraints while strictly adhering to a principle of minimal semantic intervention. To evaluate this capability, we extended the German Employment Contract Dataset with exhaustive annotations for legal validity, creating a ground truth for multi-constraint satisfaction. We benchmarked state-of-the-art LLMs, comparing emerging reasoning architectures (Gemini-2.5-pro, Claude-Opus-4.1, GPT-5, DeepSeek-v3.2-exp) against a non-reasoning baseline (GPT-4.1). Our results reveal a significant performance divergence between detection and rectification capabilities. We observe that while current reasoning models surpass the non-reasoning baseline in the rectification phase, demonstrating superior adherence to minimal intervention, they notably lag behind in classification precision during the detection phase. In addition, we observe substantial variation in rectification quality within the group of reasoning models. DeepSeek-v3.2-exp and Claude-Opus-4.1 introduce markedly fewer unnecessary semantic changes than GPT-5, which tends to rewrite clauses to a substantially greater extent despite comparable legal efficacy. Our findings suggest that practitioners should not rely on general LLM leaderboards or broad academic benchmarks when designing legal AI applications, as performance differences on specific sub-tasks can diverge significantly. Instead, effective systems require granular evaluation and, for automated contract review, a decoupling of detection and rectification to leverage the distinct strengths of specialized models.
Data: https://github.com/sebischair/Employment-Contract-Clauses-German
Published @ ICAIL 26
| Attribute | Value |
|---|---|
| Address | Singapore |
| Authors | Oliver Wardas |
| Citation | |
| Key | Wa26a |
| Research project | AI-Assisted Legal Analysis and Correction of German Employment Contracts |
| Title | LLMs for Clause Rectification in German Employment Contracts |
| Type of publication | Conference |
| Year | 2026 |
| Acronym | ICAIL |
| Project | |
| Publication URL | |
| Team members |