Back to proposals

From benchmark to regulation: Measuring what AI agents do to animals, & making it count

commitment
10 hrs/week
flexibility
Open to proposals
mentored by
Joel Christoph
Joel ChristophHarvard Kennedy School

The project

Three related directions, all building on Travel Agent Compassion (TAC) benchmark. A mentee can take one. All three are at proposal stage rather than fully specified, so a mentee with their own angle on any of them is welcome to reshape it.

  1. Extend TAC beyond travel. The benchmark measures whether an AI agent books animal-exploiting options when acting for a user who never mentions animals. Every frontier model tested scores at or below random chance. The same structure applies to procurement, meal planning, and event booking, where the scoring code carries over and the scenarios need writing from scratch. Expected output: a new domain module submitted to Inspect Evals, plus a short paper.
  2. Close the validity gaps we have named ourselves. TAC's scenario classifications are our own rather than independently validated, and there is no human baseline telling us what a travel agent would book in the same situations. A mentee could design and run expert classification with welfare and tourism researchers, or the human-agent baseline study. Expected output: a validation study that makes the benchmark defensible at a main conference.
  3. The governance route. The EU GPAI Code of Practice names risk to non-human welfare as a systemic risk but mandates no benchmark for it. There is a concrete piece of work in mapping what compliance would require, which evaluations would satisfy it, and what a provider would actually have to run. Expected output: a policy brief aimed at the AI Office and the model providers, and a submission to a governance venue.

The mentee's role

The mentee leads the work and is first author on what comes out of it. Sample tasks depending on direction:

  • Authoring and reviewing scenarios for a new domain
  • Writing the Inspect task and scorer, then submitting the PR
  • Designing the expert classification protocol and recruiting respondents
  • Running the analysis and producing the figures
  • Drafting the paper or brief
  • Mapping the Code of Practice provisions against existing evaluation practice
  • I review, unblock, and connect

Who I'm looking for

Must-haves:

  • Able to write clearly for a technical audience
  • Willing to be told a result does not hold.
  • Comfortable working async and self-directed, since I will not be checking your progress daily

Depending on direction (nobody needs all three):

  • Python and some familiarity with agentic evaluation frameworks for the benchmark work
  • Research-design and survey experience for the validation work
  • Policy-analysis experience for the governance work

Nice-to-haves:

  • Prior exposure to the EU AI Act
  • Experience getting something merged into an open-source repo
  • Background in animal welfare science

What I care about most is that you notice when the data does not say what everyone assumed it said. The most useful thing anyone did on TAC this year was read the transcripts rather than the summary statistics, and find that our headline number was partly an artifact.

Questions for applicants

  1. Which of the three directions interests you, and what would you change about it?
  2. Describe a time you found an error in your own work or someone else's. What did you do about it?
  3. What is your availability between August 31 and November 9, including any periods you will be away?

Support offered

  • Domain expertise on agentic evaluation design, measurement validity, and the EU AI governance framing
  • Detailed written feedback on drafts, which is where I am most useful
  • Shaping the research question so it answers something a lab or a regulator would act on
  • Connections into the animal welfare and AI governance networks, including the CaML and Sentient Futures people already working on TAC, and the Inspect Evals maintainers
  • Help getting work published, whether that is arXiv, a workshop, or a policy venue

A note on working style: I give fast and thorough written feedback and I am reachable async throughout the week. I would rather agree a written check-in cadence with a mentee than default to weekly calls, though I can do calls where they are the right tool.

format
Research & writing
topic
Near-term impact · Post-AGI transition · Macrostrategy
Joel Christoph

Joel Christoph

Harvard Kennedy School

I am a researcher working across AI governance, compute governance, and global public goods, based at Harvard Kennedy School. I co-developed TAC (Travel Agent Compassion), the first agentic animal welfare benchmark, now merged into the UK AI Security Institute's Inspect Evals and published at arXiv:2606.18142. My work sits where measurement meets policy: I write on the EU AI Act and the General-Purpose AI Code of Practice, which is the first major AI framework to name non-human welfare as a systemic risk. I have degrees from UCL, the Barcelona School of Economics, and the European University Institute, and I have co-founded the Equiano Institute and 10Billion.org. I am most useful to a mentee on research design, measurement validity, framing results for a policy audience, and getting work published.