What's New
You've built a Conversations AI agent — but will it handle real customer chats the way you expect? Prompt Optimizer now works with Conversations AI, so you can find out before your customers do.
It generates realistic customer scenarios, runs real chat conversations against a cloned copy of your agent, evaluates the results, pinpoints what failed, and rewrites the prompt to perform better. Your production agent stays untouched until you review and apply the improved version.
How It Works
- Configure - Build Your Test Plan
- Auto-generated scenarios: AI creates contextual test scenarios based on your agent's prompt, language, knowledge base, appointment setup, and actions — each with a customer persona, opening message, expected behaviors, and priority level.
- Reuse past scenarios: Load scenarios from earlier runs instead of starting from scratch.
- Multi-language support: Scenarios, evaluations, and optimized prompts stay aligned to your agent's language.
- Usage visibility: See daily free messages used vs. remaining, plus a pre-run checklist before you start.


- Test - Real Chats, Real Actions
- Real conversations, not simulations: Prompt Optimizer sends actual customer-style messages and waits for your agent's real replies.
- Action tracking: See when the agent triggers Appointment Booking, Human Handover, Workflow Triggers, Bot Transfer, Stop Bot, Auto Follow-up, Contact Field Updates, and Knowledge Base Queries.
- Full transcripts with AI scoring: Every chat is evaluated against expected outcomes, with clear reasoning on what passed, what failed, and why.
- Runs in the background: Leave the screen and come back to completed results.

- Improve - AI-Guided Prompt Optimization
- One-click Improvise: AI analyzes failed chats, identifies root causes, generates an improved prompt, and tests it.
- Auto Optimize: Run multiple optimization attempts automatically until your target accuracy is reached.
- Prompt diff viewer: See exactly what changed before applying anything.
- Best Variation tag: The highest-accuracy prompt is highlighted, with full history of every attempt.

Safe by Design
- Testing runs against a cloned agent - your live agent is never modified until you click "Use Prompt."
- Temporary test contacts are created for each run and cleaned up automatically, keeping your CRM free of test data.
Good to Know
- AI evaluations can vary between runs - treat accuracy scores as directional guidance.
- Real actions (bookings, workflows, handovers) execute during testing, so use test calendars and workflows where appropriate.
- Long or stuck conversations time out automatically so runs always complete.
