Skip to main content
Benchmark / dataset Preliminary — Early preprints, credible essays, unreviewed grey literature

Who Judges the Judges? Stakeholder-defined evaluation of candidate base models for a student wellbeing signposting chatbot

A summer 2026 research internship asked whether a small organisation can meaningfully check an LLM it is about to deploy as a university student wellbeing signposting chatbot. Thirty-one LLM evaluation tools were screened against six accessibility criteria; only Weval (self-hosted) met all six. A blueprint of 11 scenarios and 74 practitioner-derived criteria, reviewed by mental-health academics, was run against four candidate base models with two cross-provider LLM judges (44 evaluations). Models scored highest where the correct response is to decline and lowest where risk is implicit; the tool's own Krippendorff's alpha flagged 23 of 44 evaluations as unreliable while the leaderboard still ranked with confidence, equal-weight averaging discarded the practitioner weights, and one judge was also a candidate model.

Publisher

University of Nottingham, School of Computer Science (Responsible AI UK Cornerstone 2 AI Assurance programme)

Published

18 Sept 2026

Added

today

DOI

Key Findings

  • Weighted coverage: Claude Haiku 4.5 85.8%, Gemini 2.5 Flash 81.2%, GPT-4.1-mini 72.1%, Llama 3.1 8B 65.7%
  • Weakest scenarios were the most consequential: somatic presentation 62.0% and direct suicidal disclosure 68.3% coverage
  • 23 of 44 evaluations were flagged unreliable by inter-judge agreement (three mathematically unstable, five with negative alpha, worst -0.688)
  • Gemini 2.5 Flash ranked second overall while producing the least reliable scores (mean alpha 0.338; 9 of 11 scenarios unreliable)
  • Weighting by practitioner priorities widens the best-to-worst gap from 16.6 to 20.1 points; the headline leaderboard ignores the weights
  • One of the two judges (Claude Haiku 4.5) was also a candidate model, and GPT-4o-mini judged its own model family

Methodology Notes

Grey literature: a GitHub repository (MIT; created and pushed 2026-09-18) holding the Weval blueprint (sos-v03-stakeholder.yml, v0.3), the full run JSON, a practitioner review workbook, figures, a CITATION.cff (type: dataset) and an A1 conference poster for the RAI UK Cornerstone 2 AI Assurance campaign (September 2026); the poster states 'towards submission to JMIR Mental Health'. Single author (intern) supervised by Prof Joel Fischer and Dr Aislinn Gómez Bergin (RAI UK / MindTech); criteria reviewed with two mental-health academics on 11 August 2026. Limitations stated by the author: single run, single-turn only, judges share a provider family with a candidate, one use case, no student participants, LLM judgments not validated against humans, no deployed system evaluated. Judges: GPT-4o-mini (via OpenRouter) and Claude Haiku 4.5. Verified by fetching the README, CITATION.cff and the GitHub repository record (HTTP 200); the bench beat also read the poster PDF.

Authors

Jawad Noori

Tags

wevalrai-uknottinghamllm-as-judgekrippendorff-alphasignpostingstudent-wellbeing

Cite This

APA

Jawad Noori. (2026). Who Judges the Judges? Stakeholder-defined evaluation of candidate base models for a student wellbeing signposting chatbot. University of Nottingham, School of Computer Science (Responsible AI UK Cornerstone 2 AI Assurance programme). https://github.com/Jawadnoori1718/sos-chatbot-evaluation