Eval / eval set
A versioned collection of test inputs with expected outputs or grading rubrics, scored on every change to a prompt, model, or pipeline. The unit test suite of LLM engineering — no eval, no launch.
A versioned collection of test inputs with expected outputs or grading rubrics, scored on every change to a prompt, model, or pipeline. The unit test suite of LLM engineering — no eval, no launch.