LLM-as-judge
Using a model to grade another model's outputs at scale. Powerful but biased (position, length, self-preference) — calibrate against human labels and randomize comparisons.
Using a model to grade another model's outputs at scale. Powerful but biased (position, length, self-preference) — calibrate against human labels and randomize comparisons.