Latest AI/ML News
8160 articles · arXiv cs.AI
arXiv:2610.04515v1 Announce Type: new Abstract: Credibility is a central topic for agent-based models intended to support policy-making. Simulations m…
arXiv:2610.04510v1 Announce Type: new Abstract: Diffusion models for structured scientific generation must produce samples satisfying hard geometric c…
arXiv:2610.04475v1 Announce Type: new Abstract: Predicting cellular responses to genetic perturbations is a central capability for virtual cells and a…
arXiv:2610.04473v1 Announce Type: new Abstract: Serving an LLM means setting dozens of inference-time knobs, from per-layer KV retention to per-layer…
arXiv:2610.04472v1 Announce Type: new Abstract: Aviation accident and incident investigations generate extensive unstructured textual information cont…
arXiv:2610.04470v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to jailbreak attacks that conceal harmful intent within…
arXiv:2610.04469v1 Announce Type: new Abstract: Self-evolving vision-language models (VLMs), a form of self-improvement in which a model generates its…
arXiv:2610.04447v1 Announce Type: new Abstract: The increasing availability of generative artificial intelligence (GenAI) tools, such as ChatGPT and c…
arXiv:2610.04444v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate, complete, and transform information in…
arXiv:2610.04438v1 Announce Type: new Abstract: We present RAGrasp, a retrieval-augmented pipeline for planar parallel-jaw grasping from a compact set…
arXiv:2610.04437v1 Announce Type: new Abstract: Spreadsheet form filling requires agents to consolidate external evidence, ground values to precise ce…
arXiv:2610.04433v1 Announce Type: new Abstract: A coding agent is a language model wrapped in a harness: the system prompt, the tool set, and the cont…
arXiv:2610.04430v1 Announce Type: new Abstract: As LLMs are increasingly deployed in high-stakes professional workflows, engineers and researchers req…
arXiv:2610.04429v1 Announce Type: new Abstract: An agent that lacks the information it needs should ask rather than act, and the task definitions of a…
arXiv:2610.04418v1 Announce Type: new Abstract: The deployment of Reinforcement Learning (RL) agents in critical domains must be preceded with a pipel…
arXiv:2610.04407v1 Announce Type: new Abstract: Temporal Foundation Models (TFMs) aim to generalize across domains, datasets, and tasks. Yet, their de…
arXiv:2610.04387v1 Announce Type: new Abstract: Medical language models can produce correct diagnoses despite incomplete investigations and unsupporte…
arXiv:2610.04379v1 Announce Type: new Abstract: We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully s…
arXiv:2610.04375v1 Announce Type: new Abstract: Agents built on large language models (LLMs) build and run software through tool calls. A call reaches…
arXiv:2610.04371v1 Announce Type: new Abstract: Determining whether two programs are functionally equivalent is central to code modernization, patch v…