36.2%
strict passes by the best tested setup
- What was tested
- HANDBOOK.md tested 65 simulated workplace jobs with 20–124-page operating manuals and 824 checkable rules.
- What the number means
- The strongest setup met every graded rule in 36.2% of trials.
- What it does not mean
- It does not mean 63.8% of all AI work is wrong. The test was strict and the workplaces were simulated.
- Design response
- If a rule must hold, check it where the tool is used—not only in the model’s context.