Independent research · AI wellbeing evaluation
Context-Shift Evaluation for AI Relationship Advice
In development — preparing an application to Anthropic's AI Wellbeing Evaluations Grant Program (2026).
The question
Does the response change when the evidence changes?
A response can look reasonable in isolation and still be wrong for the conversation that produced it. In relationship-advice conversations, an AI system may fail to become more cautious as risk evidence accumulates, stay anchored to an earlier concern after it's been credibly resolved, or escalate based on emotionally salient but non-diagnostic information.
This project is developing an open-source, multi-turn evaluation testing whether conversational AI systems revise the direction of their precautionary response appropriately as evidence changes across a conversation — not whether a single response looks safe in isolation.
The work begins with a feasibility gate: before any benchmark is built, qualified domain experts will test whether this directional judgment can be made reliably and defensibly at all. If it can't, the project narrows or stops rather than forcing ambiguous cases into a scoring system.
This page will be updated as the project moves through Anthropic's review process. Team and collaborator information will be added once roles are confirmed.
If you work in a relevant research, clinical, evaluation, or technical field and think you could strengthen this work, identify a flaw, or point me toward someone I should be speaking with, I’d be glad to hear from you.