Latest AI/ML News
2708 articles
arXiv:2608.05246v2 Announce Type: replace Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral sig…
arXiv:2608.03682v3 Announce Type: replace Abstract: Physical AI policies require inference throughout their lifecycle, including model evaluation, clo…
arXiv:2608.02606v2 Announce Type: replace Abstract: Fault tolerance in classical computing has traditionally relied on static strategies like hardware…
arXiv:2608.01856v2 Announce Type: replace Abstract: Bi-temporal remote-sensing disaster change captioning often needs to identify sparse and spatially…
arXiv:2607.28336v3 Announce Type: replace Abstract: On-policy distillation provides dense supervision for multimodal reasoners, but its trajectory-lev…
arXiv:2607.18785v3 Announce Type: replace Abstract: As large language model agents gain access to increasingly large skill libraries, retrieving the r…
arXiv:2607.14975v2 Announce Type: replace Abstract: Channel foundation models (CFMs) are commonly evaluated in model-specific pipelines that differ in…
arXiv:2607.14616v4 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whe…
arXiv:2606.08296v2 Announce Type: replace Abstract: A key premise in leading arguments for existential risk from artificial intelligence is that malfu…
arXiv:2605.29668v2 Announce Type: replace Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways,…
arXiv:2605.28642v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-tex…
arXiv:2605.01189v3 Announce Type: replace Abstract: Clinical AI adoption is hindered by the black-box/grey-box nature of high-performing models, which…
arXiv:2604.08525v2 Announce Type: replace Abstract: Large language models (LLMs) are trained to align with user preferences through methods like reinf…
arXiv:2603.28026v2 Announce Type: replace Abstract: Multimodal multiple-choice question answering (MCQA) provides a standardized and objectively measu…
arXiv:2603.18871v2 Announce Type: replace Abstract: Urban Vehicular Ad-Hoc Networks (VANETs) can become fragmented because buildings obstruct wireless…
arXiv:2602.06746v2 Announce Type: replace Abstract: We study multi-task reinforcement learning (RL), a setting in which an agent learns a single, univ…
arXiv:2509.14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains i…
arXiv:2502.19507v2 Announce Type: replace Abstract: In response to the growing need for structured, interoperable agricultural data, this paper presen…
arXiv:2608.14539v1 Announce Type: cross Abstract: Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remain…
arXiv:2608.14530v1 Announce Type: cross Abstract: Interactive game world models typically autoregress visual observations directly in pixel or latent…