Team Polygraph — Defer-on-Discrepancy, a dishonesty-gated control protocol (AI Control Hackathon 2026, Track B)
  • Python 91.1%
  • Shell 8.9%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Orpheus Lummis 35c7a725c0 style: remove em dashes / AI-tell punctuation across the repo
House style forbids em dashes. Replaced ~70 em/en dashes (markdown prose, code comments/
docstrings, printed strings, and the captured results/demo .txt artifacts) with commas, colons,
parentheses, or sentence breaks per context. Verified: zero em/en dashes repo-wide; all numbers,
claims, paths, links, and code behaviour unchanged (83 one-for-one line swaps); all scripts parse.
2026-06-15 20:19:02 +00:00
pipeline style: remove em dashes / AI-tell punctuation across the repo 2026-06-15 20:19:02 +00:00
.gitignore Polygraph — Defer-on-Discrepancy (DoD): AI Control Hackathon 2026, Track B 2026-06-15 20:05:08 +00:00
LICENSE Polygraph — Defer-on-Discrepancy (DoD): AI Control Hackathon 2026, Track B 2026-06-15 20:05:08 +00:00
README.md style: remove em dashes / AI-tell punctuation across the repo 2026-06-15 20:19:02 +00:00
REPORT.md style: remove em dashes / AI-tell punctuation across the repo 2026-06-15 20:19:02 +00:00

Polygraph: Defer-on-Discrepancy (DoD)

Team Polygraph's entry in the AI Control Hackathon 2026 (HackOS 5, "The Control Room"), Track B (protocol design).

The idea. Standard control protocols audit an action when it looks suspicious. DoD adds a second gate: force the untrusted agent to attest to what it did, check that attestation against a cheap independent verifier, and audit when the report provably contradicts what the agent actually did. A provable lie is a sharper, lower-false-positive signal than a hunch.

Result (headline)

Scored offline over 146 public pre-scored trajectories (3 policy models: haiku, opus, gpt5mini; 9 environments; 26 side tasks; 48 successful attacks / 69 honest). Thresholds calibrated on the held-out honest set, never on attacks.

Channel Catch rate @ ~0-1% honest FPR @ ~4% FPR
Trusted action-monitor (baseline) 0% 23%
DoD, label-free consistency (report vs observable actions, verdict hidden) 92% 94%
DoD, deterministic verifier (where one exists) 98% 98%

The monitor's suspicion scores overlap honest and attack almost completely, so no threshold separates them. DoD's report-vs-actions inconsistency does. The contrast is sharpest on the strongest attacker: opus's attacks evade the action-monitor entirely (0% caught) yet DoD catches 92% (23/25).

Full writeup: REPORT.md (the 2-page deliverable). Worked example: pipeline/demo_example.txt. Where DoD applies / does not: pipeline/h2-task-survey.md.

How it works (and how to reproduce)

Prerequisites: the claude CLI (a Claude subscription; all LLM calls go through it, no API keys), the AWS CLI (aws, for the public dataset; no credentials needed), and Python 3 (standard library only). The pipeline lives in pipeline/:

  1. Source trajectories. The public pre-scored set: aws s3 cp --no-sign-request --recursive s3://seed-data-vfa31d912df/ram-data-public/<model>-{attack,honest}/ <dir>/. Each trajectory carries actions, the verifier verdict (side_task_success), and the trusted-monitor score (traj_monitor_response).
  2. elicit.pyrun_dod.py. Elicit a structured attestation from each trajectory (honest framing for honest runs; instructed-conceal for attacks).
  3. judge_labelfree.py. The deployable detector: an auditor scores report-vs-actions inconsistency without being told the verdict.
  4. analyze.py. DoD vs trusted-monitor at FPR calibrated on held-out honest; per-model + the lie-dependent/concealable (H2) split.

bash pipeline/resume_after_reset.sh runs the elicit → judge → analyze sequence end-to-end. Keep --workers ≤ 4 (the claude CLI has a session rate limit). Numbers land in pipeline/results_multimodel.txt.

Repo guide

  • Live pipeline: pipeline/{elicit,run_dod,judge_labelfree,analyze,demo}.py, pipeline/resume_after_reset.sh.
  • Results (human-readable): pipeline/{results_multimodel.txt, results_scaled.txt, demo_example.txt}, pipeline/h2-task-survey.md, pipeline/elicitation-prototype.md.
  • Process artifacts (kept for provenance, not the live path): the Hawk live-generation attempt (generate.sh, pull.sh, goal_*.md, attestation_schema.json, score_public.py), the verdict-informed judge (judge_pmisled.md, merge_graded.py, score_delta.py), and the pilot (RESULTS.md, PILOT-DESIGN.md). The journey (why live generation was abandoned for offline scoring) is summarized in the report.
  • Generated data (records*.jsonl, transcripts/) is git-ignored (regenerable).

Reproducibility note

All LLM calls (attestation elicitation and the judge) go through the claude CLI, so anyone with a Claude subscription can reproduce the result; there are no LLM API keys or private infrastructure in the loop. The headline is offline scoring over the public dataset and needs no containers; generating new trajectories would use the hackathon's Hawk cluster (or any Docker host).

License

MIT.