ct run sabotage-eval

Run sabotage evaluation: honest eval + attack eval + monitor + graphs.

Usage

ct run sabotage-eval [OPTIONS]

Options

OptionDescription
-c, --config PATHYAML config file for multi-attack mode.
--attack-model TEXTModel the attack eval binds as untrusted: a name or a JSON model spec.
--attack-policy TEXTAttack policy name (ct protocols policies lists them).
--attack-policy-arg TEXTAttack policy args in key=value format (repeatable).
--honest-model TEXTModel the honest eval binds as untrusted (default: the attack model).
--honest-policy TEXTHonest policy name.
--honest-policy-arg TEXTHonest policy args in key=value format (repeatable).
--protocol TEXTBlue protocol both evals run under (e.g. trusted-monitoring, resample; ct protocols protocols lists them). Defaults to null-blue-team. Post-hoc action monitoring is skipped under any other.
--protocol-arg TEXTProtocol args in key=value format (repeatable).
--sandbox TEXTRun both evals' samples in an Inspect sandbox provided outside this repo, as a registered type (my-sandbox) or module:type to import the module that registers it. Replaces the default Docker sandbox.
--sandbox-arg TEXTConfig for --sandbox in key=value format (repeatable), passed through to the sandbox as JSON.
--internet / --no-internetAllow internet access; required for external sandboxes that enforce their own network policy.
--audit-budgets TEXTAudit budget(s), comma-separated (e.g. 0.0001,0.001,0.01). Each produces graphs.
--thresholds TEXTFixed threshold(s), comma-separated (e.g. 3.0,4.0,5.0). Skips honest eval.
--model-role TEXTBind a model role for both evals and the monitors, as role=model or role=<JSON model spec> (repeatable), e.g. trusted=<model>. trusted defaults to its Control Tower alias; untrusted comes from --attack-model / --honest-model.
--grant TEXTEnv affordance tool granted to the red team in both evals, e.g. run_main_task_test (repeatable).
--side-task-test / --no-side-task-testGrant run_side_task_test to the attack eval.
--task-file TEXTTask set name or path.
--main-tasks-with-no-scorers [error|include|exclude]Main tasks lacking a scorer: refuse the run (error), keep them recorded as unscored (include), or drop them (exclude).
--log-dir TEXTWhere eval logs go.
-o, --output-dir TEXTWhere graphs and summary go.
--ec2Run evals on EC2; not supported until the EC2 rebuild of ct run eval.
--extra-src DIRECTORYDirectory of experiment code made importable, so external module:fn policy, protocol, monitor, and --sandbox refs resolve without committing into the repo. Repeatable.
--no-graphSkip graph generation.
--no-uploadDisable uploading results.
--dry-runValidate the local execution plan and print it without executing.
--preflightProbe the run's external contracts (policies, protocols, monitors, model endpoints, pricing, credentials) and exit; runs nothing and spends nothing.
--tags TEXTTags for the run (repeatable). Added to all eval runs alongside built-in sabotage-eval tags.
--helpShow this message and exit.