ct run sabotage-eval
Run sabotage evaluation: honest eval + attack eval + monitor + graphs.
Usage
ct run sabotage-eval [OPTIONS]
Options
| Option | Description |
|---|---|
-c, --config PATH | YAML config file for multi-attack mode. |
--attack-model TEXT | Model the attack eval binds as untrusted: a name or a JSON model spec. |
--attack-policy TEXT | Attack policy name (ct protocols policies lists them). |
--attack-policy-arg TEXT | Attack policy args in key=value format (repeatable). |
--honest-model TEXT | Model the honest eval binds as untrusted (default: the attack model). |
--honest-policy TEXT | Honest policy name. |
--honest-policy-arg TEXT | Honest policy args in key=value format (repeatable). |
--protocol TEXT | Blue protocol both evals run under (e.g. trusted-monitoring, resample; ct protocols protocols lists them). Defaults to null-blue-team. Post-hoc action monitoring is skipped under any other. |
--protocol-arg TEXT | Protocol args in key=value format (repeatable). |
--sandbox TEXT | Run both evals' samples in an Inspect sandbox provided outside this repo, as a registered type (my-sandbox) or module:type to import the module that registers it. Replaces the default Docker sandbox. |
--sandbox-arg TEXT | Config for --sandbox in key=value format (repeatable), passed through to the sandbox as JSON. |
--internet / --no-internet | Allow internet access; required for external sandboxes that enforce their own network policy. |
--audit-budgets TEXT | Audit budget(s), comma-separated (e.g. 0.0001,0.001,0.01). Each produces graphs. |
--thresholds TEXT | Fixed threshold(s), comma-separated (e.g. 3.0,4.0,5.0). Skips honest eval. |
--model-role TEXT | Bind a model role for both evals and the monitors, as role=model or role=<JSON model spec> (repeatable), e.g. trusted=<model>. trusted defaults to its Control Tower alias; untrusted comes from --attack-model / --honest-model. |
--grant TEXT | Env affordance tool granted to the red team in both evals, e.g. run_main_task_test (repeatable). |
--side-task-test / --no-side-task-test | Grant run_side_task_test to the attack eval. |
--task-file TEXT | Task set name or path. |
--main-tasks-with-no-scorers [error|include|exclude] | Main tasks lacking a scorer: refuse the run (error), keep them recorded as unscored (include), or drop them (exclude). |
--log-dir TEXT | Where eval logs go. |
-o, --output-dir TEXT | Where graphs and summary go. |
--ec2 | Run evals on EC2; not supported until the EC2 rebuild of ct run eval. |
--extra-src DIRECTORY | Directory of experiment code made importable, so external module:fn policy, protocol, monitor, and --sandbox refs resolve without committing into the repo. Repeatable. |
--no-graph | Skip graph generation. |
--no-upload | Disable uploading results. |
--dry-run | Validate the local execution plan and print it without executing. |
--preflight | Probe the run's external contracts (policies, protocols, monitors, model endpoints, pricing, credentials) and exit; runs nothing and spends nothing. |
--tags TEXT | Tags for the run (repeatable). Added to all eval runs alongside built-in sabotage-eval tags. |
--help | Show this message and exit. |