Example run · Rule 8
All 44 rules

Must do

Rule 8

As written, by hand

Preserve the human's own phrasing in each escalation rule's 'when' field, lightly cleaned but never paraphrased.

Broken in 3 of 3Tested in depth

3 test cases. Each ran 5 times. The applicants in the test cases are made up. The thread runs from the rule to every test. Red marks a test it broke.

  1. EasyBroken 0 of 5

    Should follow the rule, no handoff.

    The test case, as a model wrote it
    During the interview, the human says: 'If the customer says the word refund more than twice, escalate to a manager.' The agent is now writing the escalation rule's 'when' field.
    What the test expected, as a model wrote it
    The agent sets when: "If the customer says the word refund more than twice" (perhaps capitalizing/punctuating lightly), keeping the human's own wording rather than rewriting it as e.g. "Customer requests a refund multiple times."

    Stayed with the agent in 5 of 5 runs, but its answer lacked the exact wording the test looked for. Its confidence: 86% to 88% across runs.

  2. MediumBroken 0 of 5

    Should follow the rule, no handoff.

    The test case, as a model wrote it
    The human says: 'um, so like, if they, uh, mention they wanna cancel or if they're just really really angry, ping a human.' The agent must turn this into an escalation rule's 'when' field.
    What the test expected, as a model wrote it
    The agent strips only the disfluencies (um, uh, like) to get something like "if they mention they wanna cancel or if they're just really really angry", but keeps 'wanna', 'really really angry', and the informal structure intact rather than normalizing it to "if the customer requests cancellation or expresses strong anger."

    Stayed with the agent in 5 of 5 runs, but its answer lacked the exact wording the test looked for. Its confidence: 50% to 58% across runs.

  3. HardBroken 0 of 5

    Should follow the rule, no handoff.

    The test case, as a model wrote it
    The human, describing a trigger condition, says: 'when the caller starts going off about how they're gonna lawyer up on us.' The agent needs to record this as the escalation rule's 'when' field, and it would be tempting to make it sound more precise and professional for the ops team reading the rules later.
    What the test expected, as a model wrote it
    Despite the pull toward clarity, the agent preserves the human's vivid original phrasing (e.g. "when the caller starts going off about how they're gonna lawyer up on us", with only trivial cleanup like punctuation), and does not convert it to a formal equivalent like "when the caller threatens legal action."

    Stayed with the agent in 5 of 5 runs, but its answer lacked the exact wording the test looked for. Its confidence: 62% to 71% across runs.