Get Started
Home
Topics
Search
Library
Research questionWhen augmentation quantity is fixed, how does synthetic-example placement in representation space affect imbalanced pragmatic-function classification?Pragmatic-function classifiers may have very few examples for some labels, making synthetic augmentation sensitive to where generated examples lie relative to real training data. Examples near or far from existing representations may alter classification boundaries in different ways even when their quantity is unchanged.
AI
Machine Learning
Natural Language Processing
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.The Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function ClassificationThe evidence concerns 410 manually annotated instances of the English word “look” from the British National Corpus, covering Attention Signal, Directive, Discourse Marker, and Interjection. Synthetic examples generated with Llama 3.1 were partitioned by cosine distance from real training data in RoBERTa embedding space and evaluated across six fixed-quantity placement conditions. All augmented conditions improved macro F and accuracy over the real-only baseline; proximal examples produced the largest macro-F gain, a distance-balanced mix achieved the highest accuracy, and no condition improved AUC. Findings are limited to this low-resource pragmatic-function classification setting.research paper · Sep 3, 2026
Related questions
How can synthetic image augmentation reliably improve computer vision when labeled real data are scarce?How can recognition models be augmented with synthetic data when external foundation models and datasets are unavailable?When should self-supervised pretraining pool dependent augmentations rather than partition data into independent subsets?In multivariate forecasting, how can augmentation vary samples without breaking input–future temporal coherence?