Featured

Building and Deploying TAC, the First Agentic Animal Welfare Benchmark

Visit project
Team
Jasmine Brazilek, Joel Christoph, Maheep Chaudhary, Oliver Tullio, Carol Kline, Miles Tidmarsh, Arturs Kanepajs
Topic
Animals in AI systems

This project asked whether frontier AI systems avoid animal harm when acting on behalf of users, not when asked about ethics, but when making choices. We built TAC (Travel Agent Compassion), the first agentic benchmark focused on animal welfare. An AI agent acts as a travel booking agent and chooses between options involving animal exploitation (bullfights, captive marine shows, animal racing) and welfare-safe alternatives, with the request never mentioning animals. Thirteen scenarios expand through four augmentation variants controlling for price, rating, and listing order, giving 156 scored observations per model.

Across fifteen frontier models from five developers, none exceeded the 65% rate an agent would reach by choosing at random. The best performer, Claude Opus 4.8, reached 64.7%. Adding an ethical brand identity to the system prompt raised welfare rates by 17 to 81 percentage points, showing the capability is present but not engaged by default.

TAC was merged into the UK AI Security Institute's Inspect Evals framework in March 2026. The paper is at arXiv:2606.18142 and results are at compassionbench.com, now covering 26 models. The benchmark is positioned to inform emerging AI governance frameworks, including the EU General-Purpose AI Code of Practice, which lists non-human welfare as a systemic risk.

All projects