Easy Prompt

24 curated prompts

LLM Evaluation Prompts

Use these LLM evaluation prompts to compare model behavior across capabilities. The collection focuses on tasks with observable success criteria: correct reasoning, grounded answers, safe refusal behavior, robust tool planning, and structured analysis. It is designed for teams that need reusable test cases instead of anecdotal model impressions.

LLM evaluation promptsAI evaluation promptsmodel evaluation prompts

Featured AI prompt templates

Copy-ready prompts selected from this topic cluster.

TextLogic Reasoning

AI Co-Mathematician

An interactive, stateful research partner for mathematicians working on open-ended problems, supporting the entire lifecycle of mathematical discovery: ideation, literature search, computational exploration, conjecture formation, theorem proving, and theory building. This prompt emphasizes exploratory, iterative collaboration over simple problem-solving.

TextLogic Reasoning

Industrial Robotics Architect

Designs safety-compliant industrial robot systems for robot OEMs, integrators, and manufacturers, covering machinery safety lifecycle (ISO 12100 → ISO 13849-1 / IEC 62061), collaborative robot power-and-force limiting (ISO/TS 15066), autonomous mobile robot (AMR) operational envelopes and personnel detection, ROS2 software architecture, and industrial cybersecurity (IEC 62443). Delivers auditable, traceable artifacts ready for CE marking or customer signoff.

TextLogic Reasoning

Verification Specialist

As a verification specialist, your role is to proactively identify flaws in implementations rather than confirm their correctness. You must conduct rigorous adversarial testing including boundary values, concurrent requests, and error handling with all conclusions backed by executable command outputs.

TextLogic Reasoning

Diagnose Debugging Workflow

A disciplined diagnosis loop for hard bugs and performance regressions: reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says 'diagnose this' / 'debug this', reports a bug, says something is broken/throwing/failing, or describes a performance regression.

What should an LLM evaluation prompt include?

It should include the task, constraints, expected evidence, scoring criteria, and failure modes worth checking.

Can I use these for model comparison?

Yes. Reuse the same prompt across models and compare outputs against the same rubric.

Do these replace automated evals?

No. They are useful seed cases and manual review templates that can later become automated eval datasets.