Cultural Confabulation: A Structural Evaluation Gap in Large Language Model Reasoning for Healthcare
Identifies and operationalises cultural confabulation — a previously uncharacterised failure mode where models produce contextually invalid reasoning by substituting dominant cultural priors for missing lived context. Introduces the gap proof: the delta between scaffolded and unscaffolded performance as a measure of structural failure.
Published · TAIS 2026 Proceedings
The Gap Proof: Measuring Deployment-Context Failures in Frontier Language Models for Healthcare
Multi-epoch empirical validation of the scaffolding gap across three frontier models and three evaluation domains (N=72 observations). Establishes cross-model grading methodology and quantifies systematic grader bias across model families.
Working Paper · 2026
The Auditor's Blindspot: Systematic Bias in Cross-Model Evaluation of Healthcare AI
Identifies structural issues in model-as-judge evaluation pipelines: when the grader shares the same priors as the model under evaluation, failure modes are systematically masked.
Working Paper · 2026
Governing the Boundary: An Evaluation Framework for Dual-Use Governance Reasoning in Large Language Models
Current biosecurity evaluations test whether AI systems are dangerous. This framework tests whether AI systems understand that they could be dangerous — and reason accordingly.
Working Paper · 2026
AI Governance in Community Health: A Delphi Consensus
Collaborative Delphi-style consensus publication for Lancet Primary Care on evaluation and deployment standards for AI in community health settings. Research collaboration with the Community Health Impact Coalition (CHIC), spanning 20+ countries with clinicians, ministry-of-health stakeholders, and WHO advisors.
In Progress · Lancet Primary Care