The Umwelt paradigm for digital sentience
- commitment
- 10 hrs/week
- flexibility
- Open to proposals

The project
A lot of research into digital welfare first asks whether the systems have "the right stuff" (recurrent processing, global workspaces, higher-order representations) to be conscious, then get mired in intractable debates between theories of consciousness and its consciousness. I'm excited for research on a more literal reading of "what it's like": drawing a detailed picture of a system's Umwelt—environmental sensitivities, goal-seeking, adaptation, teleosemantics, the whole physical and functional picture. I think this is a promising research agenda for three reasons:
- There are good reasons to believe that experiences do not have private, intrinsic, ineffable, directly apprehensible properties. (cf. Frankish, "Illusionism as a theory of consciousness"; Dennett, "Quining Qualia"; Frankish, "Quining Diet Qualia") In that case, what most "consciousness research" is looking for does not exist, and is therefore unlikely to be fruitful. However, since experiences are still morally relevant (pain hurts!), what is therefore morally relevant is whatever remains: the functional, physical properties.
- Even if private, intrinsic, etc. properties do exist, they are tightly bound up with our Umwelten. For instance, our experiences rely on our senses, our wants come from our selective pressures, etc. So Umwelt research is likely to be fruitful even in these other worlds.
- Umwelt research is empirically tractable, whereas "consciousness research" is often (a) very confused about what it's looking for, or (b) looking for something which by definition is inaccessible to science.
I'm happy to advise both theoretical and empirical research on this agenda. Here are some ambitious project ideas (the incubator project would likely be much more small-scoped than this!):
- Example of a theoretical project in this domain: a minimum viable product for hedonium
- Even if your project is empirical, you should read the linked article to understand the underlying approach better! In the article, I apply the term "Umwelt" restrictively to only one of the five bullet points; for projects, I mean it liberally, to apply to all of the five and related subjects within the same philosophical outlook.
- Other ideas: the personal identity relation of an LLM, what is it like to be a role-player, how do autoregressive generators perceive time, etc.
- Examples of empirical projects: work along the lines of Anthropic's functional emotions, Chalmers' functional welfare, AISI's machinic psychopharmacology
Desired product looks like a detailed blog post or preprint.
The mentee's role
Empirical: Collaborating w/ me on experimental design, doing implementation yourself, then coming back to me with detailed reports of your progress & results, working in a feedback loop.
Theoretical: Lit reviews, writing arguments, mapping out cruxes & points of disagreement, amassing a great deal of evidence and forming opinions based on it. I don't expect mathematical proofs / formalization to help much at this stage, unless you see a directly useful application.
Who I'm looking for
General:
- Fluency writing a lot without using an LLM to write for you.
- Please use LLMs with literature search, red-teaming your ideas, checking your understanding, pilot testing claims, & coding implementation. But at the end of the day, you should own and understand your writing, ideas, and outputs.
- Regular documentation. You should keep a research journal shared with me, updated every work session. I'll read it before every call. If you read something, even if it turns out to be irrelevant, write it down! If you have an idea, even if you're not sure it's very good, write it down!
- Some philosophy of mind background, equivalent to taking one good seminar on the subject.
- Enough proactivity such that (a) you can identify and do things that will help the project without being explicitly told, or (b) if you don't know what to do, you contact me about where you're confused or uncertain. Don't count out option b; I don't bite, and having a clear idea of where to focus your energy can save a lot of otherwise wasted effort!
For empirical projects:
- Basic experience in one or more of the following: finetuning LLMs, logit-lens/linear probe level mech interp, creating evals; or significant experience in SWE / other computational research.
- If possible, please link some kind of legible output; a weekend of work in GitHub repo, even if the results aren't impressive, can provide significant evidence in your favor.
For theoretical projects:
- Link to some philosophical writing in your application! Doesn't need to be a journal publication. A blog post, a class paper, whatever.
Questions for applicants
- Pick one substantive claim from my project description that you think is false, underspecified, or at risk of leading to unproductive research. Present the strongest argument against it, how you would propose remedying it, and what would change your mind. (≤300 words)
- Pick one dimension of some AI system's Umwelt. Describe one experiment you could actually run in a week on an open-weight model that would tell us something about it. Say what result would surprise you. (≤400 words)
- What else are you committing to during this period, and how many hours/week?
- What do you want out of this experience?
Support offered
I have a mix of philosophical & technical background, so I'm probably most helpful at intersections: helping philosophical people w/ technical research & technical researchers with philosophy. I'm also pretty opinionated about this research agenda and what I think will be fruitful or not, so I can help with setting direction and write-ups.
- format
- Research & writing
- topic
- Artificial sentience · Macrostrategy · Longtermism

Jack Thompson
CS, philosophy, & cognitive science student @ Princeton. Fellow at Second Look Research. Alum of ARENA 8.0 & this project incubator! (Philosophy writing, website, LinkedIn)
