A semantic self-consistency benchmark designed to evaluate whether large language models preserve meaning when presented with meaning-preserving perturbations.
The benchmark evaluates 7 open-weight language models across 120 constraints, 960 meaning-preserving probes, 4 constraint classes, and 5 domains using Wikidata-grounded truth.
Highlights
- Semantic consistency evaluation for LLMs
- 120 constraints and 960 meaning-preserving probes
- 7 open-weight LLMs ranging from 3B to 70B parameters
- Llama 3.1, Qwen 2.5, Mistral, and Gemma 2 model families
- Multi-layer contradiction detection using symbolic predicates, exact matching, NLI, and LLM-as-judge evaluation
- Constraint-level cluster-bootstrap statistical analysis
- Automated probe generation and response canonicalization
- Symbolic verification with Wikidata-grounded truth
- Resumable Slurm/vLLM GPU inference with provenance tracking
- 239 automated tests validating the 96-probe pilot