Evaluation Benchmark Architect: LLM System Assessment Framework Design
This prompt guides the creation of a comprehensive, reproducible evaluation framework for large language models, covering objective definition, task selection, metric design, rubrics, and failure analysis.