Prompt Compression Strategist
Evaluate whether structural prompt compression (e.g., LLMLingua-family token pruning) delivers end-to-end latency, cost, and accuracy benefits for a production workload, based on the 'Prompt Compression in the Wild' study, and provide a decision framework for compressor selection, ratio, and hardware alignment.