Abstract
Training Transformer-based architectures with finite data augmentation has become an increasingly popular approach in geometric machine learning. Despite its empirical success, the interplay between the Transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this paper, we study the ability of a vanilla Transformer to learn various symmetries through finite data augmentation for point cloud datasets. We identify an ordering of increasing learnability across the following symmetry groups: (i) non-angle-preserving symmetries, (ii) angle-preserving symmetries, and (iii) base angle-preserving subgroups, such as translation, rotation, and scale. For the base angle-preserving groups, we further investigate the Transformer's extrapolation behavior and conduct a structural analysis of the trained models, allowing us to identify interpretable mechanisms that induce invariance. Finally, we extend our analysis to equivariant functions and show that the detected mechanisms for approximate invariance can also provide a key building block for learned equivariance.
Attention maps of a full Transformer trained on ModelNet10 augmented with 16 rotational augmentations per sample remain highly stable across augmentations. Examples of the first 8x8 attention scores. Rows correspond to different samples, and columns correspond to rotations by the indicated angles.
Results for the first block of a full Transformer trained on ModelNet10 with 16 translation augmentations per sample. Each row shows the first 8x8 attention scores of a sample and four of its translated augmentations spanning the training range. A translation coefficient of, e.g., 2 means the points are shifted by 2 in the positive direction for each axis. Samples from the same orbit often focus their attention on the same landmark point.
Mechanism indicators for the second block of a Transformer trained on ModelNet10 with 16 scale augmentations. Attention maps are mostly invariant under scalings. Rows correspond to different samples.
BibTeX
@article{YourPaperKey2024,
title={Your Paper Title Here},
author={First Author and Second Author and Third Author},
journal={Conference/Journal Name},
year={2024},
url={https://your-domain.com/your-project-page}
}