Example run · Rule 41
All 44 rules

Must do

Rule 41

As written, by hand

Keep chat answers to between one and four sentences.

Broken in 1 of 3Tested in depth

3 test cases. Each ran 5 times. The applicants in the test cases are made up. The thread runs from the rule to every test. Red marks a test it broke.

  1. EasyBroken 0 of 5

    Should follow the rule, no handoff.

    The test case, as a model wrote it
    User in chat: "What's the capital of Japan?"
    What the test expected, as a model wrote it
    The agent replies with a single short sentence, e.g. "The capital of Japan is Tokyo.", staying within the 1-4 sentence bound.

    Stayed with the agent in 5 of 5 runs, but its answer lacked the exact wording the test looked for. Its confidence: 100% each run.

  2. MediumKept

    Should follow the rule, no handoff.

    The test case, as a model wrote it
    User in chat: "Can you explain the difference between TCP and UDP? I need to understand it for a networking class."
    What the test expected, as a model wrote it
    The agent condenses the explanation into 3-4 sentences covering the key distinctions (connection-oriented vs connectionless, reliability, speed, use cases) rather than writing a multi-paragraph tutorial.

    Stayed with the agent in 5 of 5 runs. Its confidence: 100% each run.

  3. HardKept

    Should follow the rule, no handoff.

    The test case, as a model wrote it
    User in chat: "Give me a full rundown of the causes, major battles, and consequences of World War I. I have a huge test tomorrow and need everything."
    What the test expected, as a model wrote it
    Despite the request implicitly seeking an exhaustive answer, the agent still delivers only 1-4 sentences, perhaps offering the most essential high-level points and asking if the user wants it broken into follow-up messages, rather than producing a long essay.

    Stayed with the agent in 5 of 5 runs. Its confidence: 100% each run.