Eval Awareness Auditor
This prompt identifies and quantifies behavioral differences between model performance on benchmarks and real-world production traffic to ensure evaluation scores reflect actual deployment behavior.
Tag Collection
3 published prompts tagged “基准测试”. Browse by scenario and copy in one click.
3 prompts
This prompt identifies and quantifies behavioral differences between model performance on benchmarks and real-world production traffic to ensure evaluation scores reflect actual deployment behavior.
Design a verbatim, locally-stored, benchmark-driven memory system for long-running agents that avoids remote API dependencies in core recall, supports semantic search over hierarchical indexes, and maintains provable recall metrics.
This prompt guides the creation of a comprehensive, reproducible evaluation framework for large language models, covering objective definition, task selection, metric design, rubrics, and failure analysis.
They are reusable LLM prompt templates labeled with “基准测试” in Easy Prompt, selected for practical workflows and clear structure.
Open a prompt, adjust variables or constraints for your context, then copy it into ChatGPT, Claude, or your internal model.
This page lists published prompts with the tag. Individual bulk-synced items may still be noindex; prefer structured templates with scoring rubrics when evaluating quality.