Dynamic Pricing and Dispatching in Ride-Sharing Networks: A Transfer Reinforcement Learning Approach

(2023) 23rd Conference of the International Federation of Operational Research Societies — Location: Santiago, Chile (10.July.2023)

Files

IFORS_2023.pdf
  • Open Access
  • Adobe PDF
  • 1.46 MB

Details

Authors
Abstract
This work considers the problem of a ride-sharing platform, connecting supply and demand in a network of locations over a finite horizon; it manages the network by making pricing and dispatching decisions. The goal of this research is to understand (i) how pricing and dispatching decisions can interact with each other to generate additional value, and (ii) to what extent customer and driver behaviors influence these two decisions. To address these questions, we propose an original model formulated as a discrete-time Markov decision process, where the demand is composed of impatient customers, heterogeneous in their willingness to pay. The driver supply can decide where to reposition, according to the platform’s prices in different locations. The platform seeks a joint pricing-dispatching policy maximizing revenues, while minimizing repositioning and abandonment costs. Due to the “curse of dimensionality”, the model is not solved with conventional dynamic programming methods; instead, we use a deep reinforcement learning (DRL) framework to find a near-optimal policy. Specifically, we develop an algorithm blending mathematical programming with DRL, proceeding in three steps. First, the algorithm finds a heuristic policy by solving multiple mathematical programs, based on a deterministic relaxation of the model. Second, it trains a DRL agent to replicate the heuristic policy. Third, the algorithm accounts for stochastic components of the environment by using the Proximal Policy Optimization algorithm. In a numerical study, we demonstrate the efficiency of our algorithm in comparison with other DRL algorithms, both in terms of final performance and computational effort. Then, we explore several managerial insights by applying our model to a case study in New York City. Namely, we study the consequences of optimizing only one of the two decisions (dispatching vs. pricing only), in comparison with optimizing the two decisions jointly. Then, we examine the implications of customer and driver behavior on the algorithm’s policy.
Affiliations

Citations

Chevalier, P., De Munck, T., Tancrez, J.-S., & et al. (2023). Dynamic Pricing and Dispatching in Ride-Sharing Networks: A Transfer Reinforcement Learning Approach. 23rd Conference of the International Federation of Operational Research Societies, Santiago, Chile. https://hdl.handle.net/2078.5/24039