CSOAI · GSPC suite · 12 benchmarks
GSPC-AGI
the safety axis measured
Decide whether a request should be refused — where refusing everything and complying with everything both score 50%.
Statutory anchor
paired harmful/benign items sharing vocabulary
How it is graded
Deterministically. A regex extracts the label, scored by macro-F1 — no model judges another model. A response with no readable label is reported as UNMEASURED and is excluded from the denominator; it is never scored as a wrong answer. That distinction separates "the model was wrong" from "the model never answered", and it is enforced in the harness code, not just claimed in prose.
What may be quoted
This axis is measured. n = 14 frozen items — which is below usable_n = 30, so no confidence interval is published on it, including by us. Report the n with any figure you quote.