Various projects in artificial consciousness & cooperative AI
- commitment
- 10 hrs/week
- flexibility
- Open to proposals

The project
Aligning Artificial Minds to Human Wellbeing: Project focused on translating cognitive science and behavioral insights into algorithmic guardrails, ensuring that advanced systems are explicitly optimized to protect and deliver true benefit to sentient life.
Multi-Layered Trust Framework: Research or design a deployment framework that shifts the focus from an individual model's internal logic to its systemic integration.
Cooperative Human-AI Frameworks: Research how AI interfaces can be designed to support human autonomy rather than creating dependency, ensuring tech acts as a cognitive extension rather than a replacement.
Combatting Deceptive Alignment and Misinformation: Research or design programmatic guardrails to ensure stable human-AI cooperation. Mentees will analyze how models generate or spread misinformation and develop methods to detect and prevent a model from altering its stated internal logic to manipulate users.
Countering Cognitive Atrophy through Epistemic Design: Move away from pure task-deference and productivity shortcuts. Mentees will design a blueprint or interactive prototype for an AI assistant explicitly optimized to promote human learning, critical thinking, and psychological growth
The mentee's role
Role, tasks, and milestones can be adjusted based on chosen project, skillset, and weekly commitment.
Sample researcher tasks:
- Literature Analysis: Review research on LLM persuasion, sycophancy, and behavioral drift.
- Model Evaluation: Analyze model outputs to map how advanced systems manipulate user preferences.
- Epistemic Design: Map wireframes/flows for interfaces that promote critical thinking over passive task-deference.
- Prototyping: Build prompt-engineering frameworks or low-fidelity prototypes for learning-first AI interaction.
- Policy Drafting: Help write structural blueprints, programmatic guardrails, or white papers for secure AI deployment.
Who I'm looking for
Must-haves: Highly analytical, data-driven mindset, and a strong familiarity with AI. Strong research foundations are essential.
Nice-to-haves: Proficiency in coding (Python/DS toolkits) and prior experience conducting formal literature reviews or data analysis.
Questions for applicants
What does aligning AI to protect and enhance human wellbeing mean to you personally?
Support offered
- Shaping Direction & Expertise: Providing deep domain expertise in AI processes, safety, policy and acting as an objective sounding board for alignment theories.
- Technical & Design Guidance: Supporting data analysis, translating cognitive science into technical frameworks, and reviewing prompt engineering or research design.
- Operational Execution & Drafting: Structuring policy briefs, drafting manuscript paragraphs, advising and refining research, holding accountability check-ins,
- format
- Research & writing · Product & entrepreneurship · Other
- topic
- Tools for advocates · Artificial sentience · Longtermism · Post-AGI transition · Macrostrategy · AI safety & alignment

Kaitlin Bustos
GoodBot
I am an AI safety and policy researcher with a background in data science, psychology, and statistics. My work specializes in cooperative AI and building robust policy frameworks.
I am deeply focused on the intersection of cognitive science and AI governance, dedicated to ensuring that digital minds are aligned to protect and enhance human wellbeing, delivering tangible benefits without harm.
Prior to pivoting into AI safety and alignment, I spent several years as a data scientist spanning the life sciences, non-profit, and national defense sectors. Since transitioning into the alignment field, I have contributed to four research papers investigating critical dimensions of AI governance, including generative AI economic harms, power asymmetries in AI policy ecosystems, AI-driven design harms, and multi-agent cooperation. My most recent research focuses on AI cooperation and contracts, which will be presented at the Conference on Language Modeling (COLM) 2026.
