Representation choice and training-boundary composition in document stream segmentation
Abstract
This study evaluates representation choice and training-boundary composition for document stream segmentation. On Dutch OpenPSS LONG and SHORT, four document orders were compared across five training seeds while holding document exposure, boundary prevalence and pair multiplicity fixed. Term frequency–inverse document frequency (TF–IDF) features outperformed frozen multilingual MiniLM in all eight corresponding comparisons by 5.22–7.26 percentage points of boundary F1. Native ordering gave the highest LONG means. Similarity-based construction yielded a difference of −0.55 points versus random ordering in the prespecified LONG MiniLM comparison (95% bootstrap confidence interval [−1.19, +1.71]). A separate extension compared four procedures on LONG and English TABME++, using 10,000 training pairs, 3,000 validation pairs and three seeds per corpus. Selective adaptation updated the final four encoder layers for two epochs. On TABME++, it achieved 78.42% boundary F1 versus 73.72% for the matched frozen control, a difference of +4.71 points (conditional 95% cluster interval [+3.20, +6.29]). TF–IDF and frozen MiniLM with stochastic gradient descent achieved 86.43% and 85.23%, respectively. Exact document recovery and local optimization time complemented boundary scoring. The findings support comparing lexical and adapted representations under explicit training budgets and retaining observed document order as a practical baseline.
Keywords: document stream segmentation, selective fine-tuning, TF–IDF, sentence embeddings, training data composition, duplicate screening, model evaluation
Declarations
Data availability
The original datasets and pretrained model are available from the cited sources. The local research package contains source manifests, code, fixed protocols, exclusions, saved predictions and figure data.
Ethics statement
This computational study used previously released document datasets. It did not recruit participants or access private TaxDome client records.
Author contributions
Dmitriy Ulybin: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Data curation, Writing – original draft, Writing – review and editing, Visualization, and Project administration.
Funding
This research received no external funding.
Competing interests
Dmitriy Ulybin is employed by TaxDome. The author declares no competing interests related to this study.
