Teaching Stratego to Play Ball : Optimal Synthesis for Continuous Space MDPs

Legay, Axel;Jaeger, Manfred;Haahr Taankvist, Jakob;Sedwards, Sean;Larsen, Kim;et.al.
(2019) ATVA

Files

atva19.pdf
  • Open Access
  • Adobe PDF
  • 785.14 KB

Details

Authors
  • Legay, AxelUCLouvain
    Author
  • Jaeger, Manfred
    Author
  • Haahr Taankvist, Jakob
    Author
  • Sedwards, Sean
    Author
  • Larsen, Kim
    Author
Show more
Abstract
Uppaal Stratego facilitates optimization of quantitative measures on complex stochastic timed systems. In this paper we propose alternatives to the optimization algorithms of Uppaal Stratego, demonstrating that a significant improvement can be achieved in terms of convergence tendencies. In particular, we propose two online learning algorithms, using online partition refinement techniques, and argue for its theoretical convergence. We have implemented the proposed algorithms in Uppaal Stratego and support our claims with experimentson a range of models. We also provide the core algorithms as an Open Source library under the permissive LGPL license.
Affiliations

Citations

Legay, A., Jaeger, M., Haahr Taankvist, J., Sedwards, S., Jensen, P., & Larsen, K. (2019). Teaching Stratego to Play Ball : Optimal Synthesis for Continuous Space MDPs. In Manfred Jaeger, Peter Gjøl Jensen, Kim Guldstrand Larsen, Axel Legay, Sean Sedwards and Jakob Haahr Taankvist (ed.), ATVA (LNCS). Springer. https://hdl.handle.net/2078.5/226953