Reinforcement Learning for Shortest Path Problem on Stochastic Time-dependent Road Network
ID:1971 View Protection:ATTENDEE Updated Time:2021-12-03 14:43:49 Hits:251 Poster Presentation

Start Time:2021-12-17 08:44(Asia/Shanghai)

Duration:1min

Session:P2 Poster2021 » P2T1Track 1 Advanced Transportation Information and Control Engineering

Presentation File

Tips: Only the registered participant can access the file. Please sign in first.

Abstract
Finding a shortest path between two locations on stochastic time-dependent road network is an important constituent in vehicle guidance system. However, it is difficult for traditional heuristic algorithm to handle the complexity and stochasticity within the road network. In this paper, we model the stochastic time-dependent routing problem as a Markov decision process and utilize several reinforcement learning methods to solve this problem, such as Sarsa, Q-learning and Double Q-learning method. Sarsa method uses the actual Q value for iteration instead of the maximum value function used by Q-Learning, while Double Q-learning utilizes two estimators to compute the value function, which can overcome the shortcoming of overestimation. Evaluated on ten stochastic time-dependent road networks, we can draw the conclusion that Double Q-learning method outperforms other methods. Finally, the optimal paths acquired at different epochs are visualized to display the process of agent exploration.
Keywords
CICTP
Speaker
Ke Zhang
Tsinghua University

Submission Author
Ke Zhang Tsinghua University
Submit Comment
Verify Code Change Another
All Comments
Important Date
  • Conference Date

    Dec 17

    2021

    to

    Dec 20

    2021

  • Dec 16 2021

    Contribution Submission Deadline

  • Dec 24 2021

    Registration deadline

Sponsored By
Chinese Overseas Transportation Association
Chang'an University
Contact Information