<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-23T04:06:57Z</responseDate><request verb="GetRecord" identifier="oai:www.repository.cam.ac.uk:1810/386696" metadataPrefix="uketd_dc">https://api.repository.cam.ac.uk/server/oai/request</request><GetRecord><record><header><identifier>oai:www.repository.cam.ac.uk:1810/386696</identifier><datestamp>2025-07-17T01:00:50Z</datestamp><setSpec>com_1810_221783</setSpec><setSpec>com_1810_256067</setSpec><setSpec>col_1810_221784</setSpec></header><metadata><uketd_dc:uketddc xmlns:uketd_dc="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:dcterms="http://purl.org/dc/terms/" xmlns:uketdterms="http://naca.central.cranfield.ac.uk/ethos-oai/terms/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/ http://naca.central.cranfield.ac.uk/ethos-oai/2.0/uketd_dc.xsd">
   <dc:title>Enabling the discovery of novel bioactives and enzymes from billions of protein sequences</dc:title>
   <dc:identifier xsi:type="dcterms:DOI">https://doi.org/10.17863/CAM.119801</dc:identifier>
   <dc:creator>Langer, Felix</dc:creator>
   <uketdterms:advisor>Finn, Robert</uketdterms:advisor>
   <dcterms:abstract>Microorganisms are the most abundant life form on earth and are found in every naturally
occurring ecosystem. Most of these microorganisms have yet to be cultured. Metagenomics-
based methods allow the culture-independent analysis of any biome through sequencing the
whole DNA content of an environment sample and analysing it with computational methods.
MGnify is one of the most widely used platforms for the analysis of metagenomic sequences.
Since 2018, MGnify has assembled and analysed over 57,000 shotgun metagenomics datasets.
From these assemblies, over 2.4 billion non-redundant protein sequences have been identified.
This protein database outscales other protein repositories and contains a significant fraction
of functionally unexplored sequences, making it a treasure trove for protein mining. However,
due to the scale of the resource, even simply downloading the raw sequence data can be
challenging. Subsequently, search results can contain hundreds of thousands of matches,
making exploring the results difficult.
With this thesis, I outline the approaches I undertook to expand the data in the MGnify
protein database to facilitate filtering and exploration of the data. To enable this to be
accessible to all, new technical solutions were explored to allow rapid querying of the data.
Finally, to connect search results to the underlying database, I developed an interactive
platform to facilitate intuitive data mining using a comprehensive toolbox. This platform
enables the rapid identification of potentially novel bioactives and enzymes within MGnify. I
demonstrate the platform’s utility through a series of use cases focused on enzymes capable
of degrading plastics. Based on this experience and additional use cases looking at CRIPSR-
Cas systems, I extended the platform to enable the integration of multiple search results.
This extension facilitates both comparative analysis and genomic context. These additional
capabilities make it a valuable tool for discovering gene clusters and understanding the
distribution of proteins across different environments.
Overall, this thesis presents the significant expansion of one of the largest protein
databases available and the development of a versatile platform for efficient protein mining.
The ability to rapidly mine vast metagenomic datasets paves the way for discovering novel
enzymes and bioactive compounds that can be applied in industrial processes, bioremediation
or biomedical contexts.</dcterms:abstract>
   <uketdterms:institution>University of Cambridge</uketdterms:institution>
   <dcterms:issued>2024-07-31</dcterms:issued>
   <dc:type>Thesis</dc:type>
   <uketdterms:qualificationlevel>Doctoral</uketdterms:qualificationlevel>
   <uketdterms:qualificationname>Doctor of Philosophy (PhD)</uketdterms:qualificationname>
   <uketdterms:sponsor>This work was supported by EMBL and the EMBL International PhD Programme, BBSRC and Unilever</uketdterms:sponsor>
   <dcterms:isReferencedBy xsi:type="dcterms:URI">https://www.repository.cam.ac.uk/handle/1810/386696</dcterms:isReferencedBy>
   <uketdterms:embargotype>embargo</uketdterms:embargotype>
   <uketdterms:embargodate>2026-07-16</uketdterms:embargodate>
   <dc:identifier xsi:type="dcterms:URI">https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/a9bdc504-5b60-4de8-8f80-f5fa8d53dcc9/download</dc:identifier>
   <uketdterms:checksum xsi:type="uketdterms:MD5">204df9a09800229f1cb1941e34a64a74</uketdterms:checksum>
   <dcterms:license>https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e4d14855-032f-4d69-8aec-789fadf041e0/download</dcterms:license>
   <uketdterms:checksum xsi:type="uketdterms:MD5">87eda9de84448d1f82354d60eee3eb5f</uketdterms:checksum>
   <dc:rights>http://purl.org/NET/rdflicense/allrightsreserved</dc:rights>
   <dc:subject>bioinformatics</dc:subject>
   <dc:subject>computational biology</dc:subject>
   <dc:subject>protein mining</dc:subject>
</uketd_dc:uketddc>
</metadata></record></GetRecord></OAI-PMH>