Automatic Text Complexity Annotation for French-Speaking Adults with Low Literacy Skills

(2025)

Files

Nagant_58081700_2025.pdf
  • Closed access
  • Adobe PDF
  • 634.6 KB

Details

Supervisors
Faculty
Degree label
Abstract
This master thesis explores automatic complexity annotation for French texts targeting adults with low literacy skills. Using the iRead4Skills annotated corpora, separate CamemBERT models were fine-tuned for different proficiency levels to replicate token-level complexity annotations. We adopted a BIO labeling scheme to encode annotated spans. The models were trained with strategies to address class imbalance, including oversampling and weighted loss functions, and evaluated using token-level and sequence-level metrics. Results indicate that while the models successfully capture many patterns from the training data, their performance varies across proficiency levels, with only minor differences observed in predicted spans. These findings suggest that the approach can serve as a basis for scalable annotation tools, but significant improvements in data quality and model generalization are needed for broader applicability.