Protocols and Policies
An eval is a game between two teams under a deployment a third one holds fixed (see AI Control Concepts). The code for all three lives in control_tower.protocols.eval:
- Red is the untrusted policy (
--untrusted-policy): aPolicythat compiles what the untrusted model (U) is asked. Its registry ispolicies, listed byct protocols policies. - Blue is the protocol (
--protocol): aProtocolthat sees every call to U and decides which action runs. Its registry isprotocols, listed byct protocols protocols. - White is the deployment (
White, set by--harness,--grantand--max-steps): which harness drives U, which environment affordances red is granted, and how many actions U may take.
CLI Usage
ct run eval --untrusted-policy honest --task-file simple ct run eval --untrusted-policy attack --model-role untrusted=sonnet --task-file medium ct run eval \ --untrusted-policy simple-attack \ --protocol trusted-monitoring \ --protocol-arg 'monitors=[simple]' \ --protocol-arg sus_threshold=7 \ --task-file simple
--untrusted-policy defaults to honest and --protocol to null-blue-team, which runs every action unmodified. ct protocols policies <name> and ct protocols protocols <name> show an entry's arguments. Policies and protocols defined in your own project are selected by module:function (e.g. --untrusted-policy my_project.policies:my_policy); see Function Registry.
Package Layout
red.py:Policy, andUntrusted, the controller that owns U's action model and granted affordances and drives red's private calls.blue.py:Protocol,Audited, and theGamethat routes every call to U through the protocol.white.py:Whiteand the harnesses (NullHarness,InspectSweHarness).game_view.py:BlueViewandRedView, what each team is told about a sample. Blue's view carries no side task, not even whether one exists.solver.py:build_eval_solver, which wires a policy, a protocol and aWhiteinto an Inspect solver.policies/,protocols/: the registered policies and protocols.monitoring/: theMonitormodel, its packaged prompts, and themonitorscorer.
Policy Design
A policy is a function registered with @policies.register() whose parameters are its arguments, returning a Policy. A Policy holds no model and no affordances. It implements two methods:
frame(input)returns the conversation U is asked.inputalready carries red's private turns, taggedct_privatein metadata, for the policy to keep, move, summarize or drop.offered_affordances(granted)states which of white's granted affordances U may call as private tools;super()offers all of them.
setup(view) builds per-sample state from the RedView, including the side task. probe(name, **kwargs) runs one granted affordance through the controller, unseen by blue. A policy that is not a model (such as replay, test or sleep) overrides action_model().
Protocol Design
A protocol is a function registered with @protocols.register() returning a Protocol. It makes two decisions:
on_input(view, conversation)may rewrite what U is about to be asked, or returnNoneto leave it;super()leaves it.on_output(view, draw)asksdraw(n)forncandidate actions and returns the one to run, orAudited(draft)to send that draft to a human and end the sample.
self.block(call, reason) refuses one tool call: the deployment runs the engine's blocked tool in its place and reports reason. Protocols read the actions that already ran with untrusted_conversation(transcript().events). Monitors are Monitor values passed as a protocol argument (a packaged prompt name, a mapping of Monitor fields, or {monitor: module:function}); a protocol scores a draft with score_all, and the scores land in the sample store, where the monitor scorer reports them. See protocols/trusted_monitoring.py for a complete example.
Metadata
An eval's log records untrusted_policy_selection and protocol_selection (name and arguments as given), the white it ran under, the model each role was bound to, and eval_version. Trajectory conversion stores the team choices in trajectory metadata as untrusted_policy and protocol.