AMCIS 2026 · Reno, NV

Structural Biases in LLM-as-a-Judge Systems: Implications for reliable IS Adoption

Jonathan Piaget, Ulysse Rosselet, Cédric Gaspoz

University of Applied Sciences and Arts Western Switzerland (HES-SO)

August 22, 2026

Abstract

This study investigates the LLM-as-a-Judge paradigm as a critical Information System (IS) whose reliability and neutrality determine organizational adoption. Using dictionary definition evaluation as a methodologically neutral testbed, we conduct 8,000 blind pairwise comparisons between definitions from five established English dictionaries and four large language models, judged by the same four LLMs. Our results reveal a pronounced position bias (Definition A preferred 65.7% of the time), a massive self-preference bias (up to +73.3 percentage points when a model judges its own output), and moderate inter-judge agreement (Fleiss' kappa = 0.357). Lexical diversity analysis shows that LLM-generated definitions are shorter yet lexically comparable to human dictionaries. These structural biases constitute barriers to trust and adoption. Explicit bias measurement, diversified blind evaluation pipelines, and human-calibrated assessment are essential for reliable deployment of LLM-as-a-Judge systems in organizational contexts.

LLM-as-a-Judge, Algorithmic bias, AI adoption, Automated decision-making, Self-preference, Lexicography

Cite this paper

@inproceedings{piaget2026structural,
  title     = {Structural Biases in {LLM-as-a-Judge} Systems: Implications for Reliable {IS} Adoption},
  author    = {Piaget, Jonathan and Rosselet, Ulysse and Gaspoz, C{\'e}dric},
  booktitle = {Proceedings of the 32nd Americas Conference on Information Systems (AMCIS)},
  year      = {2026},
  address   = {Reno, NV, USA},
  url       = {https://aisel.aisnet.org/amcis2026/sig_svs/svs/1/}
}

Contact

Happy to talk about this work.

← Full CV