SemIntegrity: Semantic Self-Consistency Benchmark for LLMs

A benchmark for evaluating semantic consistency and meaning preservation in large language models.

A semantic self-consistency benchmark designed to evaluate whether large language models preserve meaning when presented with meaning-preserving perturbations.

The benchmark evaluates 7 open-weight language models across 120 constraints, 960 meaning-preserving probes, 4 constraint classes, and 5 domains using Wikidata-grounded truth.

Highlights

  • Semantic consistency evaluation for LLMs
  • 120 constraints and 960 meaning-preserving probes
  • 7 open-weight LLMs ranging from 3B to 70B parameters
  • Llama 3.1, Qwen 2.5, Mistral, and Gemma 2 model families
  • Multi-layer contradiction detection using symbolic predicates, exact matching, NLI, and LLM-as-judge evaluation
  • Constraint-level cluster-bootstrap statistical analysis
  • Automated probe generation and response canonicalization
  • Symbolic verification with Wikidata-grounded truth
  • Resumable Slurm/vLLM GPU inference with provenance tracking
  • 239 automated tests validating the 96-probe pilot