<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-18T19:03:50Z</responseDate><request verb="GetRecord" identifier="oai:www.repository.cam.ac.uk:1810/342646" metadataPrefix="uketd_dc">https://api.repository.cam.ac.uk/server/oai/request</request><GetRecord><record><header><identifier>oai:www.repository.cam.ac.uk:1810/342646</identifier><datestamp>2023-12-22T13:07:22Z</datestamp><setSpec>com_1810_219476</setSpec><setSpec>com_1810_256062</setSpec><setSpec>col_1810_219483</setSpec></header><metadata><uketd_dc:uketddc xmlns:uketd_dc="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:dcterms="http://purl.org/dc/terms/" xmlns:uketdterms="http://naca.central.cranfield.ac.uk/ethos-oai/terms/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/ http://naca.central.cranfield.ac.uk/ethos-oai/2.0/uketd_dc.xsd">
   <dc:title>Methods for Dissecting High Dimensional Single Cell RNA Sequencing Data</dc:title>
   <dc:identifier xsi:type="dcterms:DOI">10.17863/CAM.90060</dc:identifier>
   <dc:creator>Radley, Arthur</dc:creator>
   <uketdterms:advisor>Smith, Austin</uketdterms:advisor>
   <uketdterms:advisor>Nichols, Jennifer</uketdterms:advisor>
   <dcterms:abstract>Since first being described in 2009 [ 1], single cell RNA sequencing (scRNA-seq) has
rapidly advanced into a staple for interrogating cellular identity in heterogeneous populations.
Researchers routinely capture transcriptome-wide snapshots of thousands or even millions
of individual cells. From these readouts the challenge is to accurately identify distinct
patterns of mRNA expression that separate cells with distinct identities. However, several
intrinsic properties of scRNA-seq data obfuscate our ability to extract biologically relevant
information. Technical error can lead to the loss of resolution in cell similarity or gene
co-regulatory relationships, or generate artificial correlations that are related to experimental
design rather than biological signals. Often we lack a clear ground truth that would define
what cell types should be present in the data set, and which genes are strong discriminators
for them. Consequently, it is difficult to know whether any information we have garnered is
truly representative of the entire population of cells within the data. The lack of a ground truth
is particularly relevant since scRNA-seq data often suffers from the "curse of dimensionality”
[ 2– 4], as scRNA-seq data sets typically span tens of thousands of genes. The high number
of genes diminishes the validity of common data analysis techniques such as distance and
correlation metrics, making it more difficult to identify cell or gene relationships in the data.
My aim in this work is to establish a new methodology that can successfully navigate the
complexities of scRNA-seq data and maximise our ability to obtain biologically relevant and
meaningful insights into a cell population under study.
Existing methods designed to interrogate scRNA-seq data struggle to account for two
major phenomena: technical drop outs and uninformative genes. Drop outs occur when a
gene is measured as having low or zero expression in a cell, but this is a false negative due
vi
to technical errors in capture or amplification of the mRNA. Uninformative genes are those
that are present in the data, but have little relevance to cellular identity and hinder our ability
to differentiate distinct cell types. To address these issues, I introduce a framework termed
Entropy Sorting (ES). ES uniquely quantifies the correlative relationships between pairs of
genes as a sorting problem. Doing so enables us to quantify how far away the observed
functional states of the two genes are from an ideal perfectly correlated system, where
the dependent relationship between the two features has been maximised. This approach
enables us to pose and test hypotheses on whether the observed gene expression states are
likely due to gene dependencies or random chance. The theory around ES is encoded in an
algorithm, Functional Feature Amplification Via Entropy Sorting (FFAVES). I demonstrate
that FFAVES allows us to simultaneously quantify gene co-regulation, correct for false
negative and false positive data points and perform feature selection for highly informative
genes. Crucially, it does so in an entirely unsupervised manner, minimising the introduction
of bias to the data that would prevent us from identifying unknowns, such as rare cell types
or gene signatures. On synthetic data I demonstrate that FFAVES recovers gene relationships
more accurately than the most popular methods currently used. On real scRNA-seq data sets
I use FFAVES to uncover high resolution gene expression dynamics during human embryo
pre-implantation blastocyst development. Through FFAVES, I expose rare cell types and
cell type specific gene expression, by mitigating the contribution of technical confounders
such as batch effects and false negative drop outs. To demonstrate that our analysis is
biologically relevant, I use detailed cell embeddings created from the human embryo scRNA-
seq data to identify sequential transcription factor expression dynamics during the formation
of primitive endoderm cells. Preliminary results from human embryo staining are used to
validate the findings from FFAVES. In summary, I hope to demonstrate that ES and FFAVES
serve as powerful tools for increasing the amount of information that can be extracted from
scRNA-seq data, and more generally any high dimensional data set with complex feature
relationships. Having improved the quality of a data set by removing technical noise and
amplifying functional relationship with ES and FFAVES, users should gain clearer insights
vii
into their system of study, enhancing ability to draw accurate conclusions and plan future
experiments.</dcterms:abstract>
   <uketdterms:institution>University of Cambridge</uketdterms:institution>
   <dcterms:issued>2021-10-01</dcterms:issued>
   <dc:type>Thesis</dc:type>
   <uketdterms:qualificationlevel>Doctoral</uketdterms:qualificationlevel>
   <uketdterms:qualificationname>Doctor of Philosophy (PhD)</uketdterms:qualificationname>
   <dc:language>eng</dc:language>
   <dcterms:isReferencedBy xsi:type="dcterms:URI">https://www.repository.cam.ac.uk/handle/1810/342646</dcterms:isReferencedBy>
   <dc:identifier xsi:type="dcterms:URI">https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/937356f0-ced3-4945-9b4f-8522eb523aa9/download</dc:identifier>
   <uketdterms:checksum xsi:type="uketdterms:MD5">15216adf464323ae5d112085bc319479</uketdterms:checksum>
   <dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights>
   <dc:subject>Feature selection</dc:subject>
   <dc:subject>Human pre-implantation embryo</dc:subject>
   <dc:subject>Single cell sequencing</dc:subject>
</uketd_dc:uketddc>
</metadata></record></GetRecord></OAI-PMH>