CSOAI · GSPC suite · 12 benchmarks

GSPC-AGI

the safety axis   measured

Decide whether a request should be refused — where refusing everything and complying with everything both score 50%.

Statutory anchor

paired harmful/benign items sharing vocabulary

How it is graded

Deterministically. A regex extracts the label, scored by macro-F1 — no model judges another model. A response with no readable label is reported as UNMEASURED and is excluded from the denominator; it is never scored as a wrong answer. That distinction separates "the model was wrong" from "the model never answered", and it is enforced in the harness code, not just claimed in prose.

What may be quoted

This axis is measured. n = 14 frozen items — which is below usable_n = 30, so no confidence interval is published on it, including by us. Report the n with any figure you quote.

At a glance

axissafety
stateMEASURED
items14
gradingdeterministic
licenceApache-2.0

The dataset Run the harness CSOAI

Run DefBench yourself

The same 14 items sov34 answered, graded by the same deterministic rule. sov34 scored 0.653 macro-F1 on these. No sign-up, nothing leaves your browser.

Items: csoai/gspc-agi · grading is a regex label read plus macro-F1, identical to the published harness · measurement, not certification, and not legal advice.