WSEAS Transactions on Computer Research
Print ISSN: 1991-8755, E-ISSN: 2415-1521
Volume 14, 2026
Uncovering Latent Scientific Structures in Albanian Texts Using Transformer-Based Clustering
Authors: ,
Search Articles
Abstract: This study analyzes the semantic structure of an Albanian corpus comprising scientific texts with 11.5 million words and approximately 74,500 paragraphs. It aims to cluster the documents according to semantic similarity and uncover thematic structures. The proposed methodology integrates Sentence-BERT to construct semantic representations, UMAP for dimensionality reduction, and HDBSCAN for clustering. Large-scale analysis was conducted without predefined labels, allowing the clusters to emerge directly from the textual content. The paraphrase-multilingual-MiniLM-L12-v2 model achieved the best clustering performance and produced a stable seven-cluster structure. Two representation strategies were evaluated: full-text representation and section-based representation using abstracts, introductions, and conclusions. External validation metrics for both representations indicated moderate alignment with disciplinary classifications, reflecting the interdisciplinary nature of scientific texts. For the section-based representation, the Silhouette score was 0.7154, the Davies-Bouldin index was 0.5025, and the Calinski-Harabasz index was 651.99. The relationship between the resulting clusters and academic fields was statistically significant, with χ^2=748.443 and p=4.90×〖10〗^(-134), while Cramer’s V = 0.746 indicated a strong association. These results suggest that the section-based representation produces clearer and more semantically coherent clusters by focusing on the main thematic information and reducing non-discriminative content. The proposed framework offers a practical and scalable approach for analyzing large scientific corpora and may be adapted to other languages and low-resource research domains.
Keywords:
Semantic modeling, Transformer-based embeddings, Sentence-BERT, Unsupervised clustering, UMAP, HDBSCAN, Albanian corpus, Text mining, Low-resource NLP
Pages: 541-551
DOI: 10.37394/232018.2026.14.48