Latest AI/ML News
1452 articles · 🧠 Research
arXiv:2604.06628v2 Announce Type: replace Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while…
arXiv:2604.04074v4 Announce Type: replace Abstract: Large language model (LLM)-based reviewing systems typically assess manuscripts in isolation, leav…
arXiv:2603.17216v2 Announce Type: replace Abstract: With the advent of AI agents, automated scientific discovery is becoming an increasingly plausible…
arXiv:2603.00801v2 Announce Type: replace Abstract: Language agents increasingly act as web-enabled systems that search, browse, and synthesize inform…
arXiv:2602.12276v2 Announce Type: replace Abstract: Test-time scaling has become a standard way to improve performance and boost reliability of neural…
arXiv:2601.23049v2 Announce Type: replace Abstract: Medical calculators are fundamental to quantitative, evidence-based clinical practice. However, th…
arXiv:2511.17006v2 Announce Type: replace Abstract: Scaling test-time computation has been extended from language model reasoning to tool-augmented ag…
arXiv:2510.10813v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to domains that require reasoning about othe…
arXiv:2509.15035v2 Announce Type: replace Abstract: This study investigates the use of generative AI to support formative assessment through machine g…
arXiv:2508.09105v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) and its Multimodal Retrieval-Augmented Generation (MRAG) sign…
arXiv:2503.00992v2 Announce Type: replace Abstract: In this paper we leverage psychological methods to investigate LLMs' conceptual mastery in applyin…
arXiv:2411.15041v2 Announce Type: replace Abstract: Advanced Multimodal Large Language Models (MLLMs) struggle with recent Knowledge-based Visual Ques…
arXiv:2608.16889v1 Announce Type: cross Abstract: Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-la…
arXiv:2608.16884v1 Announce Type: cross Abstract: The current best bounds on the matrix multiplication exponent $\omega$ are obtained through a refine…
arXiv:2608.16876v1 Announce Type: cross Abstract: We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Rese…
arXiv:2608.16868v1 Announce Type: cross Abstract: A language model's output does not by itself provide verifiable evidence about the internal computat…
arXiv:2608.16844v1 Announce Type: cross Abstract: The quadratic cost of attention-based sequence models for long contexts has motivated a growing line…
arXiv:2608.16837v1 Announce Type: cross Abstract: Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet gen…
arXiv:2608.16834v1 Announce Type: cross Abstract: We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in whi…
arXiv:2608.16829v1 Announce Type: cross Abstract: Video world models approximate the stochastic distribution of physical outcomes through generative s…