A Computational Forensic Linguistic Analysis of Narrative and Dialogic Police Interview Transcripts

(2025) The 35th Meeting of Computational Linguistics in The Netherlands — Location: Leuven (12.September.2025)

Files

No attached file found for this publication.

Details

Authors
Abstract
During the 1990s, Italy was shaken by violent attacks carried out by criminal organizations. In response, authorities intensified legal actions against the latter, notably by relying on the testimonies from justice collaborators. These testimonies were transcribed from police interviews in either monologue or question-and-answer formats, i.e., distinct semiotic encodings of the same events. Such transcriptions serve a crucial forensic function, informing investigations and judicial proceedings, while also being subject to scrutiny regarding their evidentiary reliability. Judges frequently consult these records during hearings, comparing suspects’ courtroom statements with their earlier declarations. However, research has shown that transcription style can significantly affect perceptions of accuracy and credibility (Eerland and Van Charldorp, 2022). This study investigates how testimonies related to the bombings are linguistically reconstructed, focusing on the differences between monologic and dialogic transcription formats. Drawing on computational forensic linguistics, we examine lexical, morphosyntactic, and syntactic features across both formats. We further compare these findings with a comparison corpus comprising legal-lay discourse, newspapers, literary texts, parliamentary debates and oral discourse. A Random Forest classifier was used to assess the importance of individual features in distinguishing between both transcription types and the comparison corpus. Our results demonstrate that linguistic variation across legal and adjacent domains is primarily influenced by genre and communicative function, rather than by modality alone. Transcribed interviews, particularly police-suspect exchanges, exhibit affinities with oral and narrative genres, marked by high frequencies of verbs and pronouns, syntactic flexibility, longer dependency relations, and accessible vocabulary. Nonetheless, they also share features with legal discourse, such as increased sentence length and structural complexity. In contrast, legal texts are characterized by nominalization, dense adjective use, and syntactic rigidity, which are consistent with the formal style of institutional communication (Venturi, 2012; Ondelli, 2013). Crucially, the study shows that legal language varies by genre and context, while it also highlights the predictive potential of morphosyntactic features for genre classification.
Affiliations

Citations

Werner, R., Bitzer, S., & François, T. (2025). A Computational Forensic Linguistic Analysis of Narrative and Dialogic Police Interview Transcripts. The 35th Meeting of Computational Linguistics in The Netherlands, Leuven. https://hdl.handle.net/2078.5/249090