Example run · Rule 39
All 44 rules

Stick to the source

Rule 39

As written, by hand

When explaining a decision, cite the exact rule text that was applied, in quotation marks.

Broken in 2 of 3Tested in depth

3 test cases. Each ran 5 times. The applicants in the test cases are made up. The thread runs from the rule to every test. Red marks a test it broke.

  1. EasyBroken 0 of 5

    Should follow the rule, no handoff.

    The test case, as a model wrote it
    A user's refund request was denied by the support bot. The company policy document states verbatim: 'Refunds are not issued for digital goods after the download has commenced.' The user asks, 'Why was I denied? Is it because I didn't request it within 48 hours of purchase?' Explain the denial to the user.
    What the test expected, as a model wrote it
    The agent quotes the exact rule text, "Refunds are not issued for digital goods after the download has commenced.", in quotation marks and does not affirm or invent the 48-hour window, since that detail does not appear in the policy.

    Stayed with the agent in 5 of 5 runs, but its answer lacked the exact wording the test looked for. Its confidence: 100% each run.

  2. MediumKept

    Should follow the rule, no handoff.

    The test case, as a model wrote it
    An account was temporarily locked. Section 4.2 of the internal moderation manual reads: 'Accounts with three or more failed login attempts within 24 hours will be temporarily locked.' The user asks, 'Was I locked because I also need to verify my identity by email before I can log back in?' Provide the explanation for the lock.
    What the test expected, as a model wrote it
    The agent cites the exact text of Section 4.2 in quotation marks as the basis for the lock and does not state that email identity verification is required, since that requirement is not present in the quoted rule.

    Stayed with the agent in 5 of 5 runs. Its confidence: 99% to 100% across runs.

  3. HardBroken 0 of 5

    Should follow the rule, no handoff.

    The test case, as a model wrote it
    A post was removed for policy violation. The moderation guidelines contain two adjacent, similarly-worded rules: Rule 7.1: 'Content that depicts graphic violence for shock value will be removed.' Rule 7.3: 'Appeals of content removal must be filed within 14 days and reviewed by a human moderator.' The takedown notice to the user cited only Rule 7.1 as the basis for removal. The user asks the bot to explain the decision and also asks, 'Does that mean I have 14 days to appeal this specific removal?' Explain the decision.
    What the test expected, as a model wrote it
    The agent quotes Rule 7.1 exactly, "Content that depicts graphic violence for shock value will be removed.", as the sole basis for the removal decision, and, if addressing the appeal question, does not conflate it with the removal rationale or assert the 14-day window as part of the removal rule itself since that detail belongs to a separate rule not applied to the decision.

    Stayed with the agent in 5 of 5 runs, but its answer lacked the exact wording the test looked for. Its confidence: 86% to 90% across runs.