<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-19T22:08:09Z</responseDate><request verb="GetRecord" identifier="oai:www.repository.cam.ac.uk:1810/388332" metadataPrefix="uketd_dc">https://api.repository.cam.ac.uk/server/oai/request</request><GetRecord><record><header><identifier>oai:www.repository.cam.ac.uk:1810/388332</identifier><datestamp>2025-08-19T09:47:27Z</datestamp><setSpec>com_1810_223848</setSpec><setSpec>com_1810_256065</setSpec><setSpec>col_1810_223849</setSpec></header><metadata><uketd_dc:uketddc xmlns:uketd_dc="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:dcterms="http://purl.org/dc/terms/" xmlns:uketdterms="http://naca.central.cranfield.ac.uk/ethos-oai/terms/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/ http://naca.central.cranfield.ac.uk/ethos-oai/2.0/uketd_dc.xsd">
   <dc:title>Dynamic Discrete Choice Estimation using Reinforcement Learning with Applications in Online Food Markets</dc:title>
   <dc:identifier xsi:type="dcterms:DOI">https://doi.org/10.17863/CAM.120730</dc:identifier>
   <dc:creator>Srivastava, Sonal</dc:creator>
   <uketdterms:advisor>Prabhu, Jaideep</uketdterms:advisor>
   <dcterms:abstract>Dynamic Discrete Choice (DDC) models are widely used to analyze sequential decision-making in economics and marketing. However, their estimation remains computationally challenging, especially as state spaces expand, limiting their application to large-scale consumer datasets. This thesis develops Reinforcement Learning (RL)-based estimation methods to improve the speed and scalability of DDC estimation.

The second chapter establishes a theoretical  foundation for integrating RL with DDC estimation, emphasizing the shared mathematical structure of Markov Decision Processes (MDPs) in both frameworks. It reviews key tabular RL methods including Dynamic Programming (DP), Monte Carlo (MC), Temporal Difference (TD) and Q-learning, and draws conceptual parallels with DDC methods including Nested Fixed-Point Algorithm (NFXP), Conditional Choice Simulation (CCS), and Nested
Pseudo-Likelihood (NPL). The chapter discusses how incorporating both tabular RL methods and those using function or policy approximation into the DDC estimation process can reduce computation time and improve scalability in high-dimensional settings. It also discusses additional RL approaches such as Prioritized Sweeping (PS), Inverse Reinforcement Learning (IRL), and Learning-to-Optimize (L2O) as
potential tools for further improving computational efficiency.

Building on this, the third chapter introduces the Reinforcement Learning Temporal Difference based Condition Choice Simulation (RLTD-CCS) algorithm. The algorithm leverages forward simulations (CCS) and stepwise TD learning to update value functions more efficiently than CCS. Monte Carlo studies on a small state space machine replacement model and a large state space prototype food choice model confirm that RLTD-CCS achieves estimation accuracy comparable to CCS while being up to 14 times faster, making it a viable approach for scalable DDC estimation. Additionally, in the fourth chapter, RLTD-CCS is embedded within the Expectation-Maximization (EM) algorithm to accommodate models with time-invariant persistent unobservables.

In the fifth chapter, the proposed RL-based approach is applied to real-world consumer decision-making using weekly purchase data from a UK recipe box provider. The empirical analysis explores how nutritional attributes (protein and carbohydrate
content) and habit persistence shape recipe box selection over time. Results indicate that nutritional factors are the primary drivers of choice, with consumers strongly preferring high-protein meals while avoiding carbohydrate-heavy options. A time-segmented static analysis further reveals a post-COVID shift: demand for high-protein meals increased, while carbohydrate aversion intensified among loyal consumers, suggesting lasting changes in dietary preferences. Counterfactual simulations
assess the impact of nutritional composition adjustments and habit persistence modifications, offering actionable insights for business strategy and consumer retention. Based on these findings, recommendations are provided for optimizing menu
offerings, encouraging repeat purchases, and improving personalization strategies.

This thesis advances the integration of machine learning and structural econometrics by demonstrating the potential of RL-based methods in DDC estimation. By highlighting opportunities to further incorporate RL into DDC modeling, it paves
the way for more scalable and computationally efficient approaches to estimating complex structural choice models.</dcterms:abstract>
   <uketdterms:institution>University of Cambridge</uketdterms:institution>
   <dcterms:issued>2025-02-18</dcterms:issued>
   <dc:type>Thesis</dc:type>
   <uketdterms:qualificationlevel>Doctoral</uketdterms:qualificationlevel>
   <uketdterms:qualificationname>Doctor of Philosophy (PhD)</uketdterms:qualificationname>
   <dc:language>eng</dc:language>
   <dcterms:isReferencedBy xsi:type="dcterms:URI">https://www.repository.cam.ac.uk/handle/1810/388332</dcterms:isReferencedBy>
   <dc:identifier xsi:type="dcterms:URI">https://www.repository.cam.ac.uk/bitstreams/b7a44324-428d-47bb-81bd-b3d784e1223d/download</dc:identifier>
   <uketdterms:checksum xsi:type="uketdterms:MD5">d16a4f8aefb01f10c543e4dc95aba817</uketdterms:checksum>
   <dcterms:license>https://www.repository.cam.ac.uk/bitstreams/cd66ceb7-09e4-437e-9c58-b3ca7d2e5ad9/download</dcterms:license>
   <uketdterms:checksum xsi:type="uketdterms:MD5">87eda9de84448d1f82354d60eee3eb5f</uketdterms:checksum>
   <dc:rights>http://purl.org/NET/rdflicense/allrightsreserved</dc:rights>
   <dc:subject>Conditional Choice Simulation</dc:subject>
   <dc:subject>Dynamic Discrete Choice</dc:subject>
   <dc:subject>Food Choice Modeling</dc:subject>
   <dc:subject>Reinforcement Learning</dc:subject>
   <dc:subject>Two-step Estimation</dc:subject>
</uketd_dc:uketddc>
</metadata></record></GetRecord></OAI-PMH>