In an ongoing industry-university collaboration we are developing a language-parametric framework for mining code idioms in legacy systems. This modular framework has a pipeline architecture and a language-parametric meta representation of the artefacts used by each of its 5 components: source code importer, mining preprocessor, pattern miner, pattern matcher, and modernisation assistant. The pipeline enables reuse of its components across systems and languages, as well as for project partners to work on each of these components separately. An example is the exploration of novel pattern mining techniques independently of the languages on which they will be applied and the modernisation assistant in which they will be used. Our first results on mining Java and COBOL code are promising, even though challenges still lie ahead to make the framework and its constituting components truly scalable, customisable, and language independent.
Dario Di Nucci, Pham, H. S., Fabry, J., Coen De Roover, Mens, K., Molderez, T., Nijssen, S., & Zaytsev, V. (2019). A Language-Parametric Modular Framework for Mining Idiomatic Code Patterns. C E U R Workshop Proceedings, 2510(6-Dec-2019), 7. https://hdl.handle.net/2078.5/123743 (Original work published 2019)