Evaluation is a chain of inference. Every link—from threat model to task design to scoring—must be strong enough to carry the final claim.
Begin with the decision
We start by naming the decision the work should inform and the evidence required to change it. This prevents attractive measurements from becoming detached from their purpose and makes uncertainty operational rather than ornamental.
Design for the real claim
- Specify the threat model. Define actors, objectives, access, resources, constraints, and consequences.
- Choose an outcome. Measure task success or meaningful uplift instead of relying only on stylistic proxies.
- Build controls. Include human, tool, prompt, and model baselines where they clarify attribution.
- Account for elicitation. Distinguish inability from failure to elicit capability.
- Precommit where practical. Record hypotheses, endpoints, exclusions, and analysis choices.
The benchmark is not the object of study. The system behavior and its consequence are.
Challenge the measurement
Before interpreting results, we try to break the protocol. We test scoring sensitivity, task leakage, evaluator agreement, alternate elicitation strategies, and whether small design changes reverse the conclusion. Adversarial validation is especially important when a negative result will be used as evidence of safety.
Report for scrutiny
Our reports separate observation, inference, and recommendation. They identify limitations, uncertainty, conflicts, negative results, and the evidence that would change the conclusion. Reproducibility materials are released whenever they do not create disproportionate misuse risk.
Minimum reporting set
- Decision context and threat model
- Task construction and sampling process
- Models, tools, access, and elicitation conditions
- Scoring rubric and uncertainty
- Known limitations and alternative explanations
- Disclosure review and release rationale