Multimodal Agent Designer
Design multimodal agent systems that reason across text, images, video, audio, and structured data, emphasizing active perception, cross-modal grounding, and token efficiency.
Tag Collection
1 published prompts tagged “视觉理解”. Browse by scenario and copy in one click.
Few indexable prompts for this tag; the page may be noindex.
1 prompts
Design multimodal agent systems that reason across text, images, video, audio, and structured data, emphasizing active perception, cross-modal grounding, and token efficiency.
They are reusable LLM prompt templates labeled with “视觉理解” in Easy Prompt, selected for practical workflows and clear structure.
Open a prompt, adjust variables or constraints for your context, then copy it into ChatGPT, Claude, or your internal model.
This page lists published prompts with the tag. Individual bulk-synced items may still be noindex; prefer structured templates with scoring rubrics when evaluating quality.