The ability to read profoundly enriches human life. By recognizing single letters, full words, and their arrangement within phrases, we extract meaningful information from written text—an extraordinary skill made possible by the invention and evolution of writing systems. On the neural level, visual reading materializes as a complex network of areas, both visual and linguistic ones, and it involves different stages of the visual system (Dehaene et al., 2005; Yeatman & White, 2021) that culminate in a region of the brain deputed to the recognition of written words, the Visual Word Form Area (VWFA; Cohen et al., 2000, 2002; McCandliss et al., 2003) before being linked to the language network. The exact role of this area, and how it relates visual information to language, has not been completely elucidated. Given that reading commonly relies on vision, VWFA was initially conceptualized as based on the co-option of the visual processing structure that is already in place in sighted individuals (Dehaene et al., 2005; Dehaene & Cohen, 2011). In contrast, cross-modality studies state that the visual word form area may have a different, multifaceted, role in processing linguistic, and symbolic information more in general (Price & Devlin, 2003, 2011). To clarify the nature of the selectivity of VWFA, I studied the visual processing of scripts with a focus on the visual reading of Braille, a script that is not based on line junctions, the basic features that make canonical alphabets. In a first experiment (chapter 2), I investigated the reading acquisition of Braille, and compared it to a similar script based on line junctions (Line Braille). I show how the strength of the mapping between a familiar and a novel orthography drives learning more than the visual features that compose those orthographies. Through neuroimaging experiments (chapter 3), I recruited expert visual Braille readers to show that VWFA tunes to Braille, similar to how it tunes to a canonical Latin alphabet. Moreover, I present an organization of VWFA where the information of different stimuli is organized according to its linguistic content, in both Braille and Latin alphabet, but in segregated neuronal populations. Lastly, I present results from a computational approach aimed at reproducing reading in Deep Neural Networks (DNNs). I show how the synthetic model fails to reproduce the data and why this is to be attributed to the lack of linguistic knowledge. Finally, I merge the information coming from the separate studies to offer a comprehensive view of VWFA and of the influences of language and vision over reading.