Latest AI/ML News

9894 articles · arXiv cs.LG

arXiv cs.LGOct 2, 2026

arXiv:2610.00558v1 Announce Type: new Abstract: While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, thei…

arXiv cs.LGOct 2, 2026

arXiv:2610.00554v1 Announce Type: new Abstract: We evaluate hybrid quantum-classical machine learning for the reduced-order prediction of spatiotempor…

arXiv cs.LGOct 2, 2026

arXiv:2610.00545v1 Announce Type: new Abstract: We study adversarial online maximization of nonnegative, non-monotone DR-submodular functions over com…

arXiv cs.LGOct 2, 2026

arXiv:2610.00541v1 Announce Type: new Abstract: Recursive models create computational depth through parameter reuse, offering a parameter-efficient al…

arXiv cs.LGOct 2, 2026

arXiv:2610.00523v1 Announce Type: new Abstract: Conformal prediction guarantees marginal coverage, but a single calibration threshold can still spread…

arXiv cs.LGOct 2, 2026

arXiv:2610.00518v1 Announce Type: new Abstract: One-step generative models construct a static generator through iterative training-time transport. Exi…

arXiv cs.LGOct 2, 2026

arXiv:2610.00499v1 Announce Type: new Abstract: As diffusion large language models (dLLMs) become more capable, they are moving from research settings…

arXiv cs.LGOct 2, 2026

arXiv:2610.00497v1 Announce Type: new Abstract: We present Gumbel Straight Flow (GSF), a continuous flow map language model that leverages the noise-d…

arXiv cs.LGOct 2, 2026

arXiv:2610.00493v1 Announce Type: new Abstract: Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a f…

arXiv cs.LGOct 2, 2026

arXiv:2610.00446v1 Announce Type: new Abstract: As an alternative to the standard geometric analyses, we give an exact, information-theoretic analysis…

arXiv cs.LGOct 2, 2026

arXiv:2610.00445v1 Announce Type: new Abstract: Pairwise guide--transcript scores do not enforce conservation of a finite guide-loaded RISC pool when…

arXiv cs.LGOct 2, 2026

arXiv:2610.00436v1 Announce Type: new Abstract: Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Se…

arXiv cs.LGOct 2, 2026

arXiv:2610.00432v1 Announce Type: new Abstract: Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights…

arXiv cs.LGOct 2, 2026

arXiv:2610.00426v1 Announce Type: new Abstract: We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose wi…

arXiv cs.LGOct 2, 2026

arXiv:2610.00423v1 Announce Type: new Abstract: Massive activations, residual-stream coordinates with magnitudes far larger than typical activations,…

arXiv cs.LGOct 2, 2026

arXiv:2610.00418v1 Announce Type: new Abstract: Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the l…

arXiv cs.LGOct 2, 2026

arXiv:2610.00417v1 Announce Type: new Abstract: Repeated training on model-generated data can degrade later models. One possible response is to use pr…

arXiv cs.LGOct 2, 2026

arXiv:2610.00415v1 Announce Type: new Abstract: Machine-learning surrogates accelerate physical simulation, but lower prediction error need not coinci…

arXiv cs.LGOct 2, 2026

arXiv:2610.00412v1 Announce Type: new Abstract: KV cache pruning reduces long-context inference memory usage by evicting less important key-value pair…

arXiv cs.LGOct 2, 2026

arXiv:2610.00411v1 Announce Type: new Abstract: Contemporary Genomic Foundation Models (GFMs) rely on a DNA-as-a-string paradigm that employs masked t…

← PreviousPage 48 of 495Next →