<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-21T21:39:10Z</responseDate><request verb="GetRecord" identifier="oai:www.repository.cam.ac.uk:1810/252845" metadataPrefix="uketd_dc">https://api.repository.cam.ac.uk/server/oai/request</request><GetRecord><record><header><identifier>oai:www.repository.cam.ac.uk:1810/252845</identifier><datestamp>2024-06-27T10:48:29Z</datestamp><setSpec>com_1810_213747</setSpec><setSpec>com_1810_256064</setSpec><setSpec>col_1810_213748</setSpec></header><metadata><uketd_dc:uketddc xmlns:uketd_dc="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:dcterms="http://purl.org/dc/terms/" xmlns:uketdterms="http://naca.central.cranfield.ac.uk/ethos-oai/terms/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/ http://naca.central.cranfield.ac.uk/ethos-oai/2.0/uketd_dc.xsd">
   <dc:title>New approaches to modern statistical classification problems</dc:title>
   <dc:identifier xsi:type="dcterms:DOI">10.17863/CAM.16240</dc:identifier>
   <dc:creator>Cannings, Timothy Ivor</dc:creator>
   <dcterms:abstract>This thesis concerns the development and mathematical analysis of statistical procedures for&#xd;
classification problems. In supervised classification, the practitioner is presented with the&#xd;
task of assigning an object to one of two or more classes, based on a number of labelled&#xd;
observations from each class. With modern technological advances, vast amounts of data can&#xd;
be collected routinely, which creates both new challenges and opportunities for statisticians.&#xd;
After introducing the topic and reviewing the existing literature in Chapter 1, we investigate&#xd;
two of the main issues to arise in recent times.&#xd;
In Chapter 2 we introduce a very general method for high-dimensional classification,&#xd;
based on careful combination of the results of applying an arbitrary base classifier on random&#xd;
projections of the feature vectors into a lower-dimensional space. In one special case that&#xd;
we study in detail, the random projections are divided into non-overlapping blocks, and&#xd;
within each block we select the projection yielding the smallest estimate of the test error.&#xd;
Our random projection ensemble classifier then aggregates the results after applying the&#xd;
chosen projections, with a data-driven voting threshold to determine the final assignment.&#xd;
We derive bounds on the test error of a generic version of the ensemble as the number of&#xd;
projections increases. Moreover, under a low-dimensional boundary assumption, we show that&#xd;
the test error can be controlled by terms that do not depend on the original data dimension.&#xd;
The classifier is compared empirically with several other popular classifiers via an extensive&#xd;
simulation study, which reveals its excellent finite-sample performance.&#xd;
Chapter 3 focuses on the k-nearest neighbour classifier. We first derive a new global&#xd;
asymptotic expansion for its excess risk, which elucidates conditions under which the dominant &#xd;
contribution to the risk comes from the locus of points at which each class label is&#xd;
equally likely to occur, as well as situations where the dominant contribution comes from the&#xd;
tails of the marginal distribution of the features. The results motivate an improvement to the&#xd;
k-nearest neighbour classifier in semi-supervised settings. Our proposal allows k to depend&#xd;
on an estimate of the marginal density of the features based on the unlabelled training data,&#xd;
using fewer neighbours when the estimated density at the test point is small. We show that&#xd;
the proposed semi-supervised classifier achieves a better balance in terms of the asymptotic&#xd;
local bias-variance trade-off. We also demonstrate the improvement in terms of finite-sample&#xd;
performance of the tail adaptive classifier over the standard classifier via a simulation study.</dcterms:abstract>
   <uketdterms:institution>University of Cambridge</uketdterms:institution>
   <dcterms:issued>2015-11-10</dcterms:issued>
   <dc:type>Thesis</dc:type>
   <uketdterms:qualificationlevel>Doctoral</uketdterms:qualificationlevel>
   <uketdterms:qualificationname>Doctor of Philosophy (PhD)</uketdterms:qualificationname>
   <dc:language>en</dc:language>
   <dcterms:isReferencedBy xsi:type="dcterms:URI">https://www.repository.cam.ac.uk/handle/1810/252845</dcterms:isReferencedBy>
   <dc:identifier xsi:type="dcterms:URI">https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/6daaa25f-0eee-4b1b-86f7-348aec12030f/download</dc:identifier>
   <uketdterms:checksum xsi:type="uketdterms:MD5">b8a67eba5b2c9987ba4096d962b32710</uketdterms:checksum>
   <dcterms:license>https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/5141dd36-ede9-4afc-9ec7-a819ef46f607/download</dcterms:license>
   <uketdterms:checksum xsi:type="uketdterms:MD5">87eda9de84448d1f82354d60eee3eb5f</uketdterms:checksum>
</uketd_dc:uketddc>
</metadata></record></GetRecord></OAI-PMH>