Latest AI/ML News

9894 articles · arXiv cs.LG

arXiv cs.LGOct 2, 2026

arXiv:2610.01224v1 Announce Type: new Abstract: Latent world models plan by scoring candidate action sequences with distances in latent space. However…

arXiv cs.LGOct 2, 2026

arXiv:2610.01223v1 Announce Type: new Abstract: Time-series anomaly detection trades off predictive accuracy, computational efficiency, and interpreta…

arXiv cs.LGOct 2, 2026

arXiv:2610.01204v1 Announce Type: new Abstract: Autoregressive modelling has achieved remarkable success in language and sequence tasks by learning to…

arXiv cs.LGOct 2, 2026

arXiv:2610.01199v1 Announce Type: new Abstract: We consider low-budget active learning, which consists of selecting a limited number of points, the co…

arXiv cs.LGOct 2, 2026

arXiv:2610.01193v1 Announce Type: new Abstract: Counterfactual generation seeks to sample outcomes under a hypothetical intervention or decision using…

arXiv cs.LGOct 2, 2026

arXiv:2610.01181v1 Announce Type: new Abstract: We consider stochastic games with independent controlled chains and unknown transition kernels, where…

arXiv cs.LGOct 2, 2026

arXiv:2610.01175v1 Announce Type: new Abstract: Standard information bottleneck (IB) regularization constrains representations via a single scalar I(Z…

arXiv cs.LGOct 2, 2026

arXiv:2610.01173v1 Announce Type: new Abstract: The progressive growth of neural networks requires deciding when the current representation remains su…

arXiv cs.LGOct 2, 2026

arXiv:2610.01172v1 Announce Type: new Abstract: We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Mod…

arXiv cs.LGOct 2, 2026

arXiv:2610.01168v1 Announce Type: new Abstract: Time Series Anomaly Detection has received increasing attention, driven by the growing availability of…

arXiv cs.LGOct 2, 2026

arXiv:2610.01165v1 Announce Type: new Abstract: Layer interventions are widely used to probe the internal organization of language models, yet most an…

arXiv cs.LGOct 2, 2026

arXiv:2610.01153v1 Announce Type: new Abstract: Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying s…

arXiv cs.LGOct 2, 2026

arXiv:2610.01143v1 Announce Type: new Abstract: Despite strong mean accuracy, tabular foundation models (TFMs) can perform poorly on underrepresented…

arXiv cs.LGOct 2, 2026

arXiv:2610.01133v1 Announce Type: new Abstract: Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We sho…

arXiv cs.LGOct 2, 2026

arXiv:2610.01126v1 Announce Type: new Abstract: Federated learning (FL) is a communication-efficient distributed learning paradigm. However, client dr…

arXiv cs.LGOct 2, 2026

arXiv:2610.01110v1 Announce Type: new Abstract: How much can test-time computation improve a language model, and at what cost? Test-time scaling is wi…

arXiv cs.LGOct 2, 2026

arXiv:2610.01096v1 Announce Type: new Abstract: A post-hoc out-of-distribution (OOD) detector reads the activations of a trained classifier and return…

arXiv cs.LGOct 2, 2026

arXiv:2610.01076v1 Announce Type: new Abstract: Structured electronic health records (EHRs) contain a patient's clinical trajectory as a sequence of c…

arXiv cs.LGOct 2, 2026

arXiv:2610.01062v1 Announce Type: new Abstract: Activation steering provides a simple, training-free mechanism for controlling attributes of generativ…

arXiv cs.LGOct 2, 2026

arXiv:2610.01037v1 Announce Type: new Abstract: Optimizing merging coefficients for large language models can require many costly benchmark evaluation…

← PreviousPage 44 of 495Next →