Back to proposals

Two possible research tracks: (1) AI agent safeguards, (2) AI consciousness

commitment
12 hrs/week
flexibility
No mentee proposal
mentored by
Soumya Jain
Soumya JainCambridge AI Safety Hub

The project

I am proposing 2 possible project tracks and expect to pursue 1 based on applicant interest, fit and the strength of the proposed collaboration.

  1. AI incidents to agent safeguards: This project studies failures where the deployment environment turns a small model or agent error into a more serious event. The working claim is that many AI incidents begin when an ordinary-looking agent failure meets a permissive, poorly monitored, or poorly scoped deployment setup. The research will use incidents to identify those early failures, trace how they escalate, and translate the lessons into practical safeguards and evidence requirements.
  2. AI consciousness through comparative philosophy: Developing theoretical accounts of AI consciousness and sentience by relating contemporary AI consciousness frameworks to Buddhist philosophy and comparative theories of mind. The output could be a conceptual map or research paper identifying convergences, tensions, limits of analogy, and implications for AI moral status or governance.

The mentee's role

NA

Who I'm looking for

AI incidents to agent safeguards

  • Strong research, analytical, and writing skills
  • Hands-on experience with at least one aspect of agentic AI: building or deploying agents, designing evals, reviewing traces, or conducting pre- or post-deployment analysis
  • Ability to work independently, reason carefully from incomplete evidence, and contribute as a research collaborator

AI consciousness through comparative philosophy

  • Background in at least one relevant area: Buddhist philosophy, philosophy of mind, consciousness studies or comparative philosophy
  • Strong conceptual research and writing skills, including the ability to compare frameworks without flattening important differences
  • Familiarity with AI consciousness research or questions of moral status is helpful but not required

Questions for applicants

AI incidents to agent safeguards

  • Describe one agentic AI system you have built, deployed, evaluated or monitored. What did you examine, what failure pattern or risk did you identify, and what could have happened if it had gone unnoticed?
  • Select one publicly reported AI incident or near-miss and provide a link to the source. Based on the available evidence, identify the precursor failure, the deployment condition that enabled escalation, and one safeguard that could have prevented, detected, or contained it. Clearly distinguish reported facts from your own inference.

Support offered

  • Shaping the research question and scope
  • Regular feedback on analysis and writing
  • Project planning, milestones, and accountability
  • Practical insight from agentic AI deployments and enterprise AI product management
  • Practical insight from being a meditator and a Buddhist practitioner
format
Research & writing · Project scoping
topic
Other · Macrostrategy · Artificial sentience · Longtermism
Soumya Jain

Soumya Jain

Cambridge AI Safety Hub

Soumya Jain is a Research Manager at the Cambridge AI Safety Hub’s MARS V program, where she supports AI governance and technical research streams with mentors from IAPS, CNAS, Google DeepMind, LawAI, Forethought, and others. Previously, she was a MARS IV research fellow studying compute policy and AI chip export controls with Erich Grunewald at IAPS.

Before moving into AI governance, Soumya spent 5 years as a Product Manager in regulated fintech and enterprise AI, working on agentic AI deployments, enterprise context layers, and evals. She has a Bachelor’s and Master’s in Economics, and is especially interested in how complex systems behave once they leave the slide deck and meet reality. Outside work, she enjoys Buddhist philosophy, phenomenology, meditation, and hiking.