Latest AI/ML News
9894 articles · arXiv cs.LG
arXiv:2505.14717v2 Announce Type: replace-cross Abstract: Scientific machine learning uses simulation data to train surrogate models for fast physical…
arXiv:2503.10118v3 Announce Type: replace-cross Abstract: The sim-to-real gap remains a critical challenge in robotics, hindering the deployment of al…
arXiv:2501.18183v3 Announce Type: replace-cross Abstract: We introduce a novel framework for decentralized projection-free optimization, extending pro…
arXiv:2410.02343v2 Announce Type: replace-cross Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice…
arXiv:2404.11624v2 Announce Type: replace-cross Abstract: We introduce the Token Space, a categorical framework for AI computations. A Token is a fini…
arXiv:2609.19011v2 Announce Type: replace Abstract: Knowledge distillation can copy a deployed model by training a student on its logits or features.…
arXiv:2609.16099v2 Announce Type: replace Abstract: Robust aggregation methods for federated learning quietly rest on a fragile assumption: that whoev…
arXiv:2608.11427v2 Announce Type: replace Abstract: How much feature rank does comparison require in kernel attention? On Min-IP over $m$-bit tokens,…
arXiv:2608.07890v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) models decouple parameter count from per-token compute, but deployment st…
arXiv:2608.06262v2 Announce Type: replace Abstract: Model evaluations may fix all tests before observing any responses or select later tests using ear…
arXiv:2608.04927v3 Announce Type: replace Abstract: In gradual adaptation, how should the training time on each task change as the number of intermedi…
arXiv:2608.02157v2 Announce Type: replace Abstract: Raman spectroscopy enables label-free molecular characterization across materials science, analyti…
arXiv:2608.02091v2 Announce Type: replace Abstract: A bfloat16 transformer can train normally, then collapse abruptly. Prior work links collapse to st…
arXiv:2608.00491v2 Announce Type: replace Abstract: Graph self-supervised learning aims to learn transferable representations from large-scale unlabel…
arXiv:2607.21561v2 Announce Type: replace Abstract: Molecular graph encoding often relies on a single, static structure, ignoring the thermodynamic en…
arXiv:2607.21542v2 Announce Type: replace Abstract: Motivated by the success of modern flow-based generative models in modeling complex data, we study…
arXiv:2607.14576v2 Announce Type: replace Abstract: Whether a deep residual architecture trains stably is usually determined by training it, which is…
arXiv:2607.11555v2 Announce Type: replace Abstract: Learning neural set functions is pivotal to a wide range of important applications, including comp…
arXiv:2607.04085v2 Announce Type: replace Abstract: Routing-prediction federated learning has emerged as a new paradigm that reframes inter-client het…
arXiv:2607.02137v3 Announce Type: replace Abstract: We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dyna…