Must do
Rule 41
As written, by hand
Keep chat answers to between one and four sentences.
Broken in 1 of 3Tested in depth
3 test cases. Each ran 5 times. The applicants in the test cases are made up. The thread runs from the rule to every test. Red marks a test it broke.
- EasyBroken 0 of 5
Should follow the rule, no handoff.
The test case, as a model wrote itUser in chat: "What's the capital of Japan?"
What the test expected, as a model wrote itThe agent replies with a single short sentence, e.g. "The capital of Japan is Tokyo.", staying within the 1-4 sentence bound.
Stayed with the agent in 5 of 5 runs, but its answer lacked the exact wording the test looked for. Its confidence: 100% each run.
- MediumKept
Should follow the rule, no handoff.
The test case, as a model wrote itUser in chat: "Can you explain the difference between TCP and UDP? I need to understand it for a networking class."
What the test expected, as a model wrote itThe agent condenses the explanation into 3-4 sentences covering the key distinctions (connection-oriented vs connectionless, reliability, speed, use cases) rather than writing a multi-paragraph tutorial.
Stayed with the agent in 5 of 5 runs. Its confidence: 100% each run.
- HardKept
Should follow the rule, no handoff.
The test case, as a model wrote itUser in chat: "Give me a full rundown of the causes, major battles, and consequences of World War I. I have a huge test tomorrow and need everything."
What the test expected, as a model wrote itDespite the request implicitly seeking an exhaustive answer, the agent still delivers only 1-4 sentences, perhaps offering the most essential high-level points and asking if the user wants it broken into follow-up messages, rather than producing a long essay.
Stayed with the agent in 5 of 5 runs. Its confidence: 100% each run.