Introduction
The interface gives access to the prosodic analysis tools developed in the COMPEL project:
- Metrical scansion
- Enjambment detection
- Rhyme detection
- Stanza classification
How it works
The scansion module performs metrical analysis of Galician poetry, providing:
- The number of metrical syllables in each line
- The stress pattern formed by the positions of stressed metrical syllables
- The stress pattern excluding anti-rhythmic stresses (in a position consecutive with another stress)
Metrical analysis is based on the Jumper library created by Marco Remón & Gonzalo (2021), which provides scansion without prior syllabification. Jumper is specialized for Spanish, and we made some modifications to use the library with Galician texts, described in Ruiz Fabo et al. (2026a). It should be noted that Jumper is very fast, since it is not based on part-of-speech tagging and does not preprocess the text. Most of the calculation time required by scansion in GALAPAGOS is due to preprocessing the Galician text before sending it to metrical analysis.
The code is available at https://github.com/compellit/gama-sym. We also trained transformer-based neural models (encoder-decoder https://github.com/compellit/gama-trf), which are not used by GALAPAGOS because the symbolic model behaves more evenly across the different possible types of metrical licenses, as described in Ruiz Fabo et al. (2026b).
The module detects degrees of enjambment, seen as a mismatch between linguistic units and the poem's division into verse lines. The module detects the enjambment levels defined in Haider et al. (in press) and Alonso Pérez et al. (2026), which are based on Wesling's grammetrical hierarchy (1996).
| Level | Split unit | Preserved unit |
|---|---|---|
| 1 | word | phrase |
| 2 | phrase | clause |
| 3 | clause | compound sentence |
| 4 | compound sentence | |
| 5 | none (absence of enjambment) |
For each line, the module indicates its enjambment level (levels 1 to 4) with the following line, or the absence of enjambment (level 5).
Detection is implemented as line-pair classification, with a BERT model fine-tuned for the task (we created a training corpus for this). The base model used was bert-gl (Garcia, 2021).
The code, model, and training corpus will be published later.
The rhyme and stanza detection module detects consonant and assonant rhyme, as well as several types of stanza. It is a symbolic approach (rule- and dictionary-based). In recent studies (Pérez Pozo et al. 2022), this type of approach still outperforms neural models for stanza classification in Spanish.
Rhyme detection
Phonetic transcription and automatic syllabification are performed as the first step for detecting rhyme. The methods are based on Garcia and López (2011).
The module identifies the rhyme, returning its position within the rhyme scheme, its transcription in SAMPA format, and the series of vowels that are part of the rhyme (relevant in the case of assonant rhyme).
The code will be published later.
Stanza classification
The tool currently classifies the type of stanzas already delimited in the input text (with an extra line break to indicate a stanza change), but it cannot segment the text into stanzas.
Given an input where the end of each stanza has already been marked (with an extra line break), the module classifies each stanza according to an inventory of stanza forms. The rhyme scheme is computed and, based on this and on the meters of the lines, correspondences are sought with known stanza forms, as defined in Rodríguez Fer (1991).
When a match is found, the module returns the predicted stanza type for each stanza and, when possible, for the whole poem. It also indicates confidence scores that reflect the degree of correspondence with the predicted stanza forms.
The code will be published later.
Credits
The application is being developed by Xabier Suárez Cordero, Anxo Alonso Pérez, Pauline Moreau and Pablo Ruiz Fabo (PI).
The code is gradually being added to the project's GitHub group.
Metrical analysis relies on Jumper's algorithm (v. supra).
The Galician vocabulary for spelling normalization is based on a combination of terms from the dictionaries of Linguakit (Gamallo et al., 2018) and Apertium (Forcada & Tyers, 2016).
A contextual spelling normalization model was trained with texts from the corpus of the Nós project (Gamallo et al., 2024).
The phonetic transcription and automatic syllabification methods are based on Garcia and López (2011).
The work is supported by the European Union (101149659 MSCA-PF 2023).
How to cite
Suárez Cordero, X., Alonso Pérez, A., Moreau, P. & Ruiz Fabo, P. (2026). GALAPAGOS web: Interface for the Galician Automatic Poetry Anaysis System. CiTIUS - Universidade de Santiago de Compostela.
Ruiz Fabo, P., Moreau, P. & Alonso Pérez, A. (2026). Automatic Metrical Scansion of Galician Poetry: First Results. In Proceedings of PROPOR 2026. The 17th International Conference on Computational Processing of Portuguese. https://aclanthology.org/2026.propor-1.101/
References
- Alonso Pérez, Anxo, Pablo Ruiz Fabo, Thomas Haider, Pablo Rodríguez Fernández & Pablo Gamallo (2026). O GalAPAgoS (Galician Automatic Poetry Analysis System) Comeza a Camiñar. III Xeira CLARIAH-GAL. Santiago de Compostela. doi: 10.5281/zenodo.20554754.
- Carballo Calero, Ricardo (1966). Gramática elemental del gallego común. Vigo: Galaxia.
- Forcada, Mikel L. & Tyers, Francis M. (2016). Apertium: a free/open source platform for machine translation and basic language technology. In Proceedings of the 19th Annual Conference of the European Association for Machine Translation: Projects/Products. Riga, Latvia.
- Freixeiro Mato, Xosé (2006). Gramática da lingua galega I - Fonética e fonoloxía. Vigo: Edicións A Nosa Terra.
- Gamallo, Pablo, Marcos Garcia, César Piñeiro, Rodrigo Martínez-Castaño and Juan C. Pichel (2018). LinguaKit: a Big Data-based multilingual tool for linguistic analysis and information extraction. In Fifth Conference on Social Network Analysis, Management and Security, pp. 239-244. Available at IEEE Xplore.
- Gamallo, P., Rodríguez, P., Paniagua, S., Bardanca, D., Pichel, J. R., & Garcia, M. (2024). Open Generative Large Language Models for Galician. Procesamiento del Lenguaje Natural, 73, pp. 259-270. Available at SEPLN.
- Garcia, Marcos & Isaac González López (2011). Conversión fonética automática con información fonológica para el gallego. Procesamiento del Lenguaje Natural, 47, pp. 283-291.
- Haider, Thomas, Pablo Ruiz Fabo, Timo Baumann & Clara I. Martínez Cantón (2026, accepted). Enjambement as Syntactic Disruption in Historical and Contemporary German and Spanish Poetry.
- Marco Remón, G., & Gonzalo, J. (2021). Escansión automática de poesía española sin silabación. Procesamiento del Lenguaje Natural, 66, pp. 77-87. Available at SEPLN.
- Pérez Pozo, Álvaro, Javier de la Rosa, Salvador Ros, Elena González-Blanco, Laura Hernández & Mirella de Sisto (2022). A Bridge Too Far for Artificial Intelligence?: Automatic Classification of Stanzas in Spanish Poetry. Journal of the Association for Information Science and Technology, 73(2), pp. 258-267. doi: 10.1002/asi.24532.
- Rodríguez Fer, Claudio (1991). Arte literaria. Vigo: Xerais.
- Ruiz Fabo, Pablo, Pauline Moreau & Anxo Alonso Pérez (2026a). Automatic Metrical Scansion of Galician Poetry: First Results. In Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pp. 994-1004. Salvador, Brazil: Association for Computational Linguistics. Available at ACL Anthology.
- Ruiz Fabo, Pablo, Anxo Alonso Pérez, Pablo Rodríguez Fernández & Pablo Gamallo (2026b). Automatic Metrical Scansion of Poetry in a Low-Resource Setting. In LLMs4SSH @ LREC 2026: Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities. Palma de Mallorca. doi: 10.5281/zenodo.19701641.
- Wesling, Donald (1996). The scissors of meter: grammetrics and reading. University of Michigan Press.