Latest AI/ML News

268 articles · arXiv cs.AI

arXiv cs.AIAug 17, 2026

arXiv:2608.13598v1 Announce Type: new Abstract: Agent evaluation relies almost entirely on outcome metrics such as success rate, which capture whether…

arXiv cs.AIAug 17, 2026

arXiv:2608.13591v1 Announce Type: new Abstract: High-confidence errors in large language models are often treated as evidence of fragile internal infe…

arXiv cs.AIAug 17, 2026

arXiv:2608.13577v1 Announce Type: new Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman au…

arXiv cs.AIAug 17, 2026

arXiv:2608.13574v1 Announce Type: new Abstract: LLM agents increasingly operate as execution systems that invoke tools, modify local state, use persis…

arXiv cs.AIAug 17, 2026

arXiv:2608.13573v1 Announce Type: new Abstract: Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are esse…

arXiv cs.AIAug 17, 2026

arXiv:2608.13567v1 Announce Type: new Abstract: The human brain exhibits a striking degree of functional specialization, with distinct networks suppor…

arXiv cs.AIAug 17, 2026

arXiv:2608.13565v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures scale large language models (LLMs) while preserving computation…

arXiv cs.AIAug 17, 2026

arXiv:2608.13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automat…

← PreviousPage 14 of 14