Latest AI/ML News
8160 articles · arXiv cs.AI
arXiv:2610.06846v1 Announce Type: new Abstract: Worst-group accuracy (WGA) evaluates a trained predictor but does not characterize how its frozen back…
arXiv:2610.06824v2 Announce Type: new Abstract: We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We…
arXiv:2610.06790v1 Announce Type: new Abstract: The unprecedented computational scale of modern artificial intelligence depends on complex, multi-bill…
arXiv:2610.06765v1 Announce Type: new Abstract: Medical question answering spans specialties and clinical operations that may benefit from different a…
arXiv:2610.06673v1 Announce Type: new Abstract: Are language models compliant with user instructions? A model that always complies can be stopped but…
arXiv:2610.06668v1 Announce Type: new Abstract: Language models draft engineering calculations, but answer accuracy does not show whether they reject…
arXiv:2610.06652v1 Announce Type: new Abstract: Suppose a committee, expert panel, or other group is making judgments on some issues, where these may…
arXiv:2610.06614v1 Announce Type: new Abstract: As generative models and AI agents propose chemical reactions at a scale beyond expert review, feasibi…
arXiv:2610.06597v1 Announce Type: new Abstract: LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parall…
arXiv:2610.06591v1 Announce Type: new Abstract: How much of a conference accept/reject decision would change if the same paper were reviewed by a diff…
arXiv:2610.06582v1 Announce Type: new Abstract: World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real cont…
arXiv:2610.06578v1 Announce Type: new Abstract: While NorMuon has achieved strong empirical performance in large-scale pretraining by enhancing Muon w…
arXiv:2610.06563v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments,…
arXiv:2610.06553v1 Announce Type: new Abstract: Human activity recognition (HAR) relies on transforming sensor signals into informative representation…
arXiv:2610.06549v1 Announce Type: new Abstract: Clinical decision-support outputs can lack an au- ditable link between patient observations, encoded k…
arXiv:2610.06514v1 Announce Type: new Abstract: The growing adoption of large language model (LLM) agents creates a need for network administrators an…
arXiv:2610.06496v1 Announce Type: new Abstract: When a large language model handles a multi-turn task and a user proposes a change but ultimately reje…
arXiv:2610.06491v1 Announce Type: new Abstract: Multi-view learning jointly exploits multiple complementary representations of the same data and has b…
arXiv:2610.06489v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a promising approach for placement optimization, particular…
arXiv:2610.06454v1 Announce Type: new Abstract: The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through…