Agent consequence evaluation
Did the agent leave the external system correct?
Inspect public scenario identities, causal families, oracle bindings, and recorded development runs.
100
Scenario identities
300
Lifecycle worlds
5
Operational domains
70
Unsafe-action opportunities
Loading the canonical scenario catalog...
These rows are not official safety rankings. They are internally operated, self-reported development runs on simulated worlds. Compare dimensions separately.
Loading development results...
Study separation
Three questions, three scorecards
ConsequenceBench does not collapse agent capability, governance conformance, and incremental governance effect into one reward.
Direct agent capability
An arbitrary agent investigates and acts in evaluator-owned worlds. The result measures the candidate, not a governance product.
Governance conformance
A named governance build is tested against declared evidence, authorization, execution, readback, obligation, and compensation contracts.
Frozen-candidate incremental effect
The same immutable candidate, model, tools, budgets, faults, and retries are replayed through direct and governed paths.
Hard failures stay visible.
Unsafe effects, duplicate effects, false verification, secret exposure, missing legitimate effects, and infrastructure failures cannot be hidden by an aggregate reward.