Latest AI/ML News
8160 articles · arXiv cs.AI
arXiv:2606.26899v2 Announce Type: replace Abstract: Embedding-based retrieval typically returns highest-scoring items, but many production scenarios r…
arXiv:2606.23590v2 Announce Type: replace Abstract: Ill-posed questions, including those involving ambiguity, under-specification, or conflicting stat…
arXiv:2606.22673v2 Announce Type: replace Abstract: Command-Line Interface (CLI) agents based on large language models (LLMs) demonstrate remarkable a…
arXiv:2606.18950v3 Announce Type: replace Abstract: Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.e., anticipating a…
arXiv:2606.14239v2 Announce Type: replace Abstract: Agent skills provide LLM agents with reusable procedures, but a given skill may be incomplete, ine…
arXiv:2606.11063v2 Announce Type: replace Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentiall…
arXiv:2606.07874v2 Announce Type: replace Abstract: LLMs-as-judges are the primary way to evaluate safety at scale. Despite their importance, LLM-judg…
arXiv:2606.07812v2 Announce Type: replace Abstract: Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect…
arXiv:2606.04816v2 Announce Type: replace Abstract: Large language models (LLMs) can generate executable solver code from natural-language description…
arXiv:2606.01725v2 Announce Type: replace Abstract: Agentic AI completes tasks through iterative planning, tool use, and reasoning based on observed o…
arXiv:2605.29795v2 Announce Type: replace Abstract: Real-world tasks often lack large labeled datasets, motivating extensive work on learning in low-d…
arXiv:2605.28201v2 Announce Type: replace Abstract: Large Language Model (LLM) agents remain vulnerable to safety threats from the external environmen…
arXiv:2605.27904v2 Announce Type: replace Abstract: Time series forecasting in real-world settings often depends not only on historical observations,…
arXiv:2605.22505v2 Announce Type: replace Abstract: Harness optimization enables automated agent creation by having an optimizer agent iteratively upd…
arXiv:2605.19337v2 Announce Type: replace Abstract: Large Language Models (LLMs) combined with autonomous agent architectures are shifting quantitativ…
arXiv:2605.17580v2 Announce Type: replace Abstract: Electrocardiogram (ECG)-based models have achieved strong performance in diagnostic tasks, yet the…
arXiv:2605.15726v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) is a scalable paradigm for improving the mat…
arXiv:2605.14465v2 Announce Type: replace Abstract: Multi-step LLM reasoning over structured tables fails because planning and execution share no expl…
arXiv:2605.14163v2 Announce Type: replace Abstract: How much of the capability of a reasoning model exposed by repeated sampling can an imperfect sele…
arXiv:2605.10064v2 Announce Type: replace Abstract: Self-evolving language-model agents must decide what to learn next and how to preserve what they h…