← Back to all articles
arXiv cs.LGOctober 2, 2026

Attention Kernels for Learning Maps Between Heavy-Tailed Measures

Excerpt

arXiv:2610.00564v1 Announce Type: new Abstract: Operator learning on probability measures can be accomplished with transformers. For measures with polynomial tails, the exponential weighting in softmax can make the corresponding measure-level attention integrals diverge. This motivates replacing the exponential with slower-growing functions. We construct two benchmarks for operator learning on measures with closed-form targets. We use these benchmarks to study attention kernel growth and data tr