Codifying the Judge: Scalable Evaluation via Program Distillation
arXiv:2607.22561v1 Announce Type: new Abstract: LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer
延伸阅读
相关资讯
SF-AMS: Strategic Forgetting for Structured Memory in LLM Agent
今天MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models
今天DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
今天Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
今天