Flag as unhelpful
One tap. Adds to the review queue. Good when the full correction can happen later.
As AI models got dramatically better, the whole premise of HelloRep shifted. Merchants no longer needed to write conversation logic. They just needed to correct the AI when it was wrong.


The Flow Builder asked merchants to think like programmers,define every case, anticipate every branch, connect every outcome. That made sense when the AI needed explicit instructions to function. But as language models improved, the AI started handling open-ended conversations correctly without being told what to do. It could infer intent, handle variations of the same question, and respond in a brand's tone without explicit rules for each scenario.
What it still needed was domain knowledge and correction. The AI didn't know the merchant's return policy or their specific product details. And when it got something wrong, the only way to fix it was to go back into the flow builder and add another branch,which often introduced new complexity and new problems.
The better mental model was already familiar. When you hire a support agent, you don't write a decision tree for them. You tell them about the store, you answer their questions, and when they handle a case badly you tell them what they should have done instead. Test & Train is that process applied to AI.
A merchant opens the Test & Train screen and has a conversation with the AI as if they were a customer. They ask the questions their real customers ask. If the AI answers correctly, they move on. If it doesn't, they correct it right there,suggesting a better answer, adding a fact to the knowledge base, or adjusting the AI's tone. Those corrections accumulate into an AI that's calibrated to their specific store.

The redesigned canvas. Node types are colour-coded and visually distinct. Branching logic is grouped and labelled. The overall structure reads left to right,evaluate intent, decide, act, resolve.
The testing part was straightforward to design. The correction part wasn't. When a merchant sees a wrong answer, what exactly are they correcting? The AI's understanding of the question? The wording of the answer? An underlying factual gap? The correction interface needed to handle all of these without making the merchant choose between them explicitly.
The answer was a tiered model. A merchant can flag a response as unhelpful with one tap. They can suggest a replacement answer directly in the thread. Or they can add an FAQ entry that addresses the gap the wrong answer exposed. Each path feeds back into the AI differently,direct corrections affect specific responses quickly, FAQ additions improve the broader knowledge base, and flagged responses feed into ongoing calibration.
One tap. Adds to the review queue. Good when the full correction can happen later.
Replace the AI's response with what it should have said. Directly shapes that specific case.
Turn a gap in the AI's knowledge into an FAQ entry. Fixes the category, not just the instance.
Change tone, depth, or scope at the brand level,affects all future responses, not just one.
One thing that made Test & Train genuinely useful in practice was that we didn't make merchants guess what to test. The interface surfaces suggested questions based on the merchant's product category and platform,the things customers of similar stores ask most often. A fashion store gets questions about sizing, returns, and material. A supplements store gets questions about ingredients, dosage, and shipping. This seeded testing sessions and helped merchants find gaps they wouldn't have thought to look for.

The Conversations view feeds directly into Test & Train. Merchants can find real conversations where the AI underperformed, pull them into the training flow, and turn live failures into corrections.
Test & Train didn't replace the Flow Builder,it changed its role. The builder remained useful for structured, predictable workflows: a returns process that always follows the same steps, an upsell triggered at a specific moment, a handover that fires when a customer mentions a particular phrase. These cases still benefit from explicit logic. But the majority of conversations,open-ended questions, browsing, product discovery,are now handled by the AI directly, shaped through training rather than programming.
Test & Train shifted HelloRep from something that required technical thinking to something that required product knowledge. Any merchant who knows their store can train the AI. That's a fundamentally different audience than the one the flow builder could ever reach.