Finish PhD in AI Safety/ AI Auditing (Self-Funded) (Computer Science at UoM)
Finish self-funded AI auditing research PhD in +5 months. Need to rent GPUs for experiments and present at conferences. (OPTIONAL: I need to pay for living costs)
Finish self-funded AI auditing research PhD in +5 months. Need to rent GPUs for experiments and present at conferences. (OPTIONAL: I need to pay for living costs)
Project Details
Updated 07/28/26 · Edited by orgI aim to get 2-3 papers (with 1 already mostly done and proven to work) out in these 5 months and to present these papers at 2 conferences. I am a 3rd year PhD in Computer Science at The University of Manchester and these papers will conclude my PhD thesis, which needs to be submitted by the end of December 2026.
This late in my degree, I have shifted from upskilling to actively contributing to meaningful AI safety research, meaning that my output is now more-so bottlenecked by access to resources rather than skill. I believe my research will be more impactful if I am able to run larger-scale experiments (more experiments and on models closer to the frontier). Additionally, research is more seen and taken more seriously when it passes conference peer review, but conferences usually mandate that the author attends to present their work, requiring entry cost and travel funding.
My current research is to do with improving automated auditing methods of LLMs. Specifically, my work extends the Anthropic-built and UK-AISI-used BLOOM tool, improving its elicitation of rare misalignment behaviours with fewer samples, enabling faster and better pre-deployment model auditing. I have produced a draft of the paper that contains all the sections completed except for the results, which currently sit as placeholders.I have already completed the headline experiments (BLOOM vs new BLOOM, on 9 behaviours and 4 models) so just have some ablations/ analysis, as well as comparison methods remaining to do.
OPTIONAL:
Additionally, since the degree is self-funded, I have been paying for tuition and my living expenses using my own savings, partly gotten from intermittent University work and funding. I budgeted incorrectly for my rent so I think it will run out by November, but I am working part time as a BlueDot programme facilitator (4 hours a week for 5 weeks) plus doing teaching assistant work (a one-off 30 hour contract) and this should hopefully cover the gap, so it's only the compute/ conferences that needs funding. In the past I have also volunteered to facilitate AI Safety Camp and I currently still volunteer for PauseAI local group organising, but this only takes 1-4 hours a week and my PhD is still my full-time profession.
Theory of Impact
Updated 07/28/26 · By grantmaking.aiBoth my current work and future work aim to decrease existential risk from unaligned AI with immediate impacts, focusing on outputs that are immediately usable. For example, I am already running the final experiments for my current line of work which contributes to solving the following problems: (1) manual auditing of models is becoming harder as there are more models being released more quickly, with more numerous high-expertise behaviours to check for; (2) since deployed models get used at such a high scale, long-tail risky behaviours that were not found during auditing are likely to appear at deployment.
My solution involves automating auditing and using certain techniques like importance sampling to elicit rare behaviours without needing to do a million samples, allowing developers to catch alignment failures prior to deployment, reducing chance of misaligned AI taking actions in the real world without developer oversight. Additionally, my work is likely to move onto the challenge of (3) auditing in spite of model evaluation awareness and alignment faking.
OPTIONAL:
Additionally, this work will enable me to graduate and go on to do further AI safety research within high-skill fellowships (e.g. MATS) or jobs at safety organisations, as well as continue to volunteer towards field-building and civic engagement/ PauseAI related work.
People
Updated 07/28/26 · Edited by orgResearcher
Funding Details
- -
- Dec 31, 2026
- 5 months (though PhDs often have extensions)
- -
- -
- -
- -
- -
- -
- -
Track Record
I have been involved with 10 research papers that have passed workshop peer-review, (incomplete list here: https://scholar.google.com/citations?hl=en&user=9Vf0mWkAAAAJ&view_op=list_works&sortby=pubdate). 1 paper, for which I am the first author rather than a collaborator or supervisor, also got accepted into a large conference (ACL Findings 2026). This paper was to do with using linear probes to monitor the internals of LLMs for misaligned behaviours and was completed during a 3 month interruption from my PhD, as part of the competitive and funded AI safety fellowship at LASR Labs. I extended this work by then part-time supervising 12 researchers over 5 months as part of the voluntary AI Safety Camp programme, which led to 3 further workshop papers being produced. I have been a participant of AI Safety Camp before, myself, which involved work on mechanistic interpretability. For my PhD, I currently have 1 paper being reviewed for a medium conference (EMNLP 2026). For my current PhD work, I have already done the MVP experiments and the results are reasonably promising that I have a good contribution here. In terms of my safety credentials, I have been an EA for the last 6 years (with this being my first year of the 10% giving pledge) and got into Computer Science/ AI research because of 80 000 hours. I have organised for the EA student society, ran AI Safety reading groups and am currently running the Manchester branch of PauseAI.
Discussion
No comments yet. Be the first to share your thoughts.