Uppaal Stratego facilitates optimization of quantitative measures on complex stochastic timed systems. In this paper we propose alternatives to the optimization algorithms of Uppaal Stratego, demonstrating that a significant improvement can be achieved in terms of convergence tendencies. In particular, we propose two online learning algorithms, using online partition refinement techniques, and argue for its theoretical convergence. We have implemented the proposed algorithms in Uppaal Stratego and support our claims with experimentson a range of models. We also provide the core algorithms as an Open Source library under the permissive LGPL license.
Legay, A., Jaeger, M., Haahr Taankvist, J., Sedwards, S., Jensen, P., & Larsen, K. (2019). Teaching Stratego to Play Ball : Optimal Synthesis for Continuous Space MDPs. In Manfred Jaeger, Peter Gjøl Jensen, Kim Guldstrand Larsen, Axel Legay, Sean Sedwards and Jakob Haahr Taankvist (ed.), ATVA (LNCS). Springer. https://hdl.handle.net/2078.5/226953