Latest AI/ML News
2784 articles · arXiv cs.CL
arXiv:2405.15454v5 Announce Type: replace Abstract: The prevalence of Large Language Models (LLMs) in critical applications highlights the need for co…
arXiv:2609.28449v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason abou…
arXiv:2609.28442v1 Announce Type: cross Abstract: Reordering a set of mathematical rules without changing its meaning should preserve the correct answ…
arXiv:2609.28344v1 Announce Type: cross Abstract: Audio-language models (ALMs) integrate acoustic perception with the knowledge encoded in language mo…
arXiv:2609.28274v1 Announce Type: cross Abstract: The final safeguard against rogue AI behavior is the human ability to shut systems down. It has been…
arXiv:2609.28212v1 Announce Type: cross Abstract: Language modeling using Transformers has become commonplace despite their fixed computational depth…
arXiv:2609.28197v1 Announce Type: cross Abstract: As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring…
arXiv:2609.28053v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) training requires global load balance to prevent expert under-utilization a…
arXiv:2609.28029v1 Announce Type: cross Abstract: The key-value (KV) cache of autoregressive transformers can be viewed as a fourth-order tensor spann…
arXiv:2609.27844v1 Announce Type: cross Abstract: Claim denial management costs U.S. healthcare approximately $260 billion annually in administrative…
arXiv:2609.27822v1 Announce Type: cross Abstract: A common multi-agent design asks agents to report confidence and lets the highest-scoring agent spea…
arXiv:2609.27811v1 Announce Type: cross Abstract: Online platforms have become arenas for the public contestation of climate change, shaping how scien…
arXiv:2609.27756v1 Announce Type: cross Abstract: Large language models are increasingly asked to analyze data and report what the results mean, a tas…
arXiv:2609.27749v1 Announce Type: cross Abstract: The surge in AI-assisted generation of educational materials has outpaced our capacity to validate t…
arXiv:2609.27657v1 Announce Type: cross Abstract: Solutions based on large language models (LLMs) often rely on temperature sampling to improve accura…
arXiv:2609.27581v1 Announce Type: cross Abstract: Step Law gives power-law formulas for the optimal peak learning rate eta* and batch size B* when pre…
arXiv:2609.27532v1 Announce Type: cross Abstract: Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool call…
arXiv:2609.27470v1 Announce Type: cross Abstract: Recent video-language models increasingly adopt hybrid architectures that interleave linear and full…
arXiv:2609.27408v1 Announce Type: cross Abstract: Benchmarks for vision-language models offer their answer choices in some convention: a letter, a col…
arXiv:2609.27382v1 Announce Type: cross Abstract: Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers,…