Agent Eval Designer
Design real-world useful AI agent evaluations that separate model capability, harness quality, tool reliability, and environment noise through executable tasks, safety boundaries, and multi-dimensional scoring.
Tag Collection
1 published prompts tagged “安全评估”. Browse by scenario and copy in one click.
Few indexable prompts for this tag; the page may be noindex.
1 prompts
Design real-world useful AI agent evaluations that separate model capability, harness quality, tool reliability, and environment noise through executable tasks, safety boundaries, and multi-dimensional scoring.
They are reusable LLM prompt templates labeled with “安全评估” in Easy Prompt, selected for practical workflows and clear structure.
Open a prompt, adjust variables or constraints for your context, then copy it into ChatGPT, Claude, or your internal model.
This page lists published prompts with the tag. Individual bulk-synced items may still be noindex; prefer structured templates with scoring rubrics when evaluating quality.