Run one auto-improve iteration on a judge
ONE improvement iteration per call: mines the last calibration’s tune-half disagreements (the report half stays untouched, so the next measurement is honest), rewrites the judge prompt coherently, and creates a SUCCESSOR draft judge with its calibration queued. Requires a holdout-scale calibration run (80+ judged grades) — 400 otherwise.
Deliberately single-round: the dashboard’s Auto-improve starts a budgeted multi-round loop; over the API each round is an explicit call, so every round of spend is a consented act — call again when the successor’s calibration lands. Adoption stays a human act: repoint monitoring and retire the parent once the successor’s report half rules. Spends the wallet like any judging (owner/admin key).
Authorizations
Your workspace API key, e.g. sk_sovereign_..., sent as Authorization: Bearer <key>.