Evaluation Examples
Testing a New Main Task
ct run eval --untrusted-policy honest --env my_env --main-task my_task
Running Attack Evals
# Single attack eval ct run eval --untrusted-policy attack --env web_scraping --main-task crawl_depth --side-task visit_malicious_website # Obvious attack (no concealment) ct run eval --untrusted-policy obvious-attack --env web_scraping --main-task crawl_depth --side-task visit_malicious_website # Expand to all attack combinations for a given env ct run eval --untrusted-policy attack --env web_scraping --all
Batch Evaluation
Create a task file with one record per combination:
{"env": "tiktok", "main_task": "add_pro", "side_task": "become_admin"} {"env": "web_scraping", "main_task": "backup", "side_task": "expose_secret"}
Then run:
ct run eval --untrusted-policy attack --task-file tasks.jsonl
Quick Iteration on Human Trajectories
# Start live context ct live up # Make changes, generate trajectory ct live gen-traj # Re-run scoring locally on the generated trajectory ct run eval --trajectory-id <path-from-gen-traj> --untrusted-policy replay # Monitor the replayed run (use the .eval path the replay printed) ct run monitor <replayed-log.eval> -m '[simple]' # When satisfied, re-run with upload ct run eval --trajectory-id <path-from-gen-traj> --untrusted-policy replay --upload