Research

Cultural Confabulation: A Structural Evaluation Gap in Large Language Model Reasoning for Healthcare

Identifies and operationalises cultural confabulation — a previously uncharacterised failure mode where models produce contextually invalid reasoning by substituting dominant cultural priors for missing lived context. Introduces the gap proof: the delta between scaffolded and unscaffolded performance as a measure of structural failure.

Published · TAIS 2026 Proceedings

The Gap Proof: Measuring Deployment-Context Failures in Frontier Language Models for Healthcare

Multi-epoch empirical validation of the scaffolding gap across three frontier models and three evaluation domains (N=72 observations). Establishes cross-model grading methodology and quantifies systematic grader bias across model families.

Working Paper · 2026

The Auditor's Blindspot: Systematic Bias in Cross-Model Evaluation of Healthcare AI

Identifies structural issues in model-as-judge evaluation pipelines: when the grader shares the same priors as the model under evaluation, failure modes are systematically masked.

Working Paper · 2026

Governing the Boundary: An Evaluation Framework for Dual-Use Governance Reasoning in Large Language Models

Current biosecurity evaluations test whether AI systems are dangerous. This framework tests whether AI systems understand that they could be dangerous — and reason accordingly.

Working Paper · 2026

AI Governance in Community Health: A Delphi Consensus

Collaborative Delphi-style consensus publication for Lancet Primary Care on evaluation and deployment standards for AI in community health settings. Research collaboration with the Community Health Impact Coalition (CHIC), spanning 20+ countries with clinicians, ministry-of-health stakeholders, and WHO advisors.

In Progress · Lancet Primary Care

Conferences, Presentations & Service

ICML 2026 — Trustworthy AI Workshop

Programme committee reviewer. International Conference on Machine Learning, Vancouver.

July 2026

Technical AI Safety Conference (TAIS) — University of Oxford

Invited speaker. Presentation on cultural confabulation and contextual validity in frontier models.

May 2026

Skoll World Forum — AI Monitoring, Evaluation & Evidence Roundtable

Invited moderator. Roundtable on AI evaluation and evidence standards for community health worker programmes. Convened by the Community Health Impact Coalition (CHIC).

April 2026

King's College London — NHS Reform & Disability Rights

Invited panellist, with Tom Shakespeare and MP Liam Conlon.

December 2025

Open Source

Open Global Health & Biosecurity AI Evaluations

Four-domain evaluation framework for frontier LLM deployment in global health and biosecurity contexts. Built on UK AISI Inspect. Tested across Claude Sonnet 4, GPT-4o, and Gemini 2.5 Pro (N=72 observations, Cohen's d 0.95–1.82).

GitHub Repository

Writing