AI Engineering Research Note2026Canonical long-form →

LLM-as-Judge Harness

LLM-as-judge pipeline validated against human raters : Cohen's κ = 0.58 with bootstrap CI and position-bias measured.

AI EngineeringPublished Mon Jan 12 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

EVAL-FIRST31 GATESNDA-CLEANPUBLIC DATAALPHASIGNAL > NOISEREPRODUCIBLENOTEBOOK-COMMITTEDCITABLEBIBTEX + DOI

Cite this research note

A copy-pasteable BibTeX entry. The canonical long-form with figures, code, and full methodology lives at the URL in the entry. please link there, not here, when citing in a paper or a thread.

@techreport{macion2026ai03judgeharness,
  author       = {Macion, Christian T.},
  title        = {LLM-as-Judge Harness},
  institution  = {Independent research},
  year         = {2026},
  date         = {2026-01-12},
  note         = {Public-data reproducible. Canonical long-form: https://christianmacion-portfolio.pages.dev/projects/ai/03-judge-harness/},
  url          = {https://christianmacion-portfolio.pages.dev/projects/ai/03-judge-harness/}
}

Canonical surface

The full research note. methodology, code, gates run, and reproduced metrics. lives on the projects index.

  • Canonical URL/projects/ai/03-judge-harness ↗
  • Headline metric(s)
    • 0.58Cohen's κ (vs human)
    • 0.60 ± 0.02Pass rate
    • 17%Position bias
  • Tags
    • llm-as-judge
    • evaluation
    • cohen-kappa
    • bias
    • validation

Read the full research note.

The canonical long-form on /projects walks through the methodology, evaluation gates, code, and reproducibility manifest.