<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-21T05:13:16Z</responseDate><request verb="GetRecord" identifier="oai:www.repository.cam.ac.uk:1810/345671" metadataPrefix="uketd_dc">https://api.repository.cam.ac.uk/server/oai/request</request><GetRecord><record><header><identifier>oai:www.repository.cam.ac.uk:1810/345671</identifier><datestamp>2023-12-22T13:55:41Z</datestamp><setSpec>com_1810_219481</setSpec><setSpec>com_1810_256065</setSpec><setSpec>col_1810_219482</setSpec></header><metadata><uketd_dc:uketddc xmlns:uketd_dc="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:dcterms="http://purl.org/dc/terms/" xmlns:uketdterms="http://naca.central.cranfield.ac.uk/ethos-oai/terms/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/ http://naca.central.cranfield.ac.uk/ethos-oai/2.0/uketd_dc.xsd">
   <dc:title>A causal perspective on model robustness: case studies in health and sensor data</dc:title>
   <dc:identifier xsi:type="dcterms:DOI">10.17863/CAM.93093</dc:identifier>
   <dc:creator>Hasthanasombat, Apinan</dc:creator>
   <uketdterms:advisor>Mascolo, Cecilia</uketdterms:advisor>
   <dcterms:abstract>Robustness of predictive deep models is a challenging problem with many implications.
It is of particular importance when models are used in safety-critical applications, 
such as healthcare. However, there is yet to be agreement on a comprehensive definition 
on what it means for a model to be robust, and a theory on why these issues arise.
Given the general nature of the problem, existing work related to robustness is spread 
across different areas of research. Existing research has considered a range of robustness 
aspects, for instance robustness to small input perturbations, which arise from the study 
of adversarial examples, but there is also robustness to different domains for the same 
task, and robustness issues which arise from object placement, transplanting, lighting, 
weather conditions, or object style, as some examples.

This thesis explores a formulation of robustness in terms of the assumed structural causal 
model (SCM) which generates the observed data.The SCM allows these different types of 
robustness issues to be viewed in a unifying way. Using this view, this work furthers 
the connection between prediction robustness and the assumed structural causal model by 
suggesting that optimising for prediction performance across a diverse set of distributions
from the same SCM will move the model closer to the causal predictor of the target variable,
providing a theoretical foundation to optimise purely for prediction in the setting where 
training and testing data are not independently and identically distributed.


Formulating robustness in this way suggests that large deep models should, in general, 
be more susceptible to robustness issues; while some of these issues have been observed 
in applications such as computer vision, it has been less discussed in others. We 
investigate the robustness of state-of-the-art deep (SotA) classifiers in human activity 
recognition using a new proposed benchmark informed by the causal formulation, and show 
that a simpler model is at least as robust as SotA deep models whilst being at least ten 
times faster to train. The causal view of robustness additionally hints at the idea that 
less data can be beneficial for robustness, contrary to popular belief that more data 
is always better. To test this idea, a data selection algorithm is proposed based on 
inverting the idea of a popular causal inference procedure for tabular data. The robustness 
of a model trained on the selected subset of data is evaluated through synthetic and 
semi-synthetic data experiments. Under certain conditions the data subset improves 
robustness and subsequently data efficiency.</dcterms:abstract>
   <uketdterms:institution>University of Cambridge</uketdterms:institution>
   <dcterms:issued>2022-09-01</dcterms:issued>
   <dc:type>Thesis</dc:type>
   <uketdterms:qualificationlevel>Doctoral</uketdterms:qualificationlevel>
   <uketdterms:qualificationname>Doctor of Philosophy (PhD)</uketdterms:qualificationname>
   <dc:language>eng</dc:language>
   <uketdterms:sponsor>Cambridge Trust and King's College</uketdterms:sponsor>
   <dcterms:isReferencedBy xsi:type="dcterms:URI">https://www.repository.cam.ac.uk/handle/1810/345671</dcterms:isReferencedBy>
   <dc:identifier xsi:type="dcterms:URI">https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/7429e762-89cb-4108-bcf6-6065ea1f23f7/download</dc:identifier>
   <uketdterms:checksum xsi:type="uketdterms:MD5">ac92e4f157d499d97624ad4fd62aa154</uketdterms:checksum>
   <dc:rights>https://www.rioxx.net/licenses/all-rights-reserved/</dc:rights>
   <dc:subject>causality, model robustness</dc:subject>
</uketd_dc:uketddc>
</metadata></record></GetRecord></OAI-PMH>