Biography
James Oldfield is a Postdoctoral Research Assistant in the Department of Engineering Science at the University of Oxford, working on AI safety and interpretability for language models. During his PhD at Queen Mary University of London, his work developed scalable methods to decompose machine learning models into human-interpretable components, with the broader goal of making AI systems more transparent and aligned with human values. Since 2023, his work has appeared annually at leading venues (ICLR, NeurIPS, and ICML). He has also received top reviewer awards each year from 2024 to 2026.
Research Interests
AI safety
Interpretability
Current Research Projects
Building better AI safeguards
Monitoring model behavior
Understanding model internals, and interfaces for doing so