Latest AI/ML News
2784 articles · arXiv cs.CL
arXiv:2601.02031v3 Announce Type: replace-cross Abstract: Pretraining of large language models is not only expensive but also prone to certain trainin…
arXiv:2412.04236v2 Announce Type: replace-cross Abstract: Data-driven approaches to philosophy have emerged as a valuable tool for studying the histor…
arXiv:2609.01068v2 Announce Type: replace Abstract: The heavy-tailed distribution of output lengths in Large Language Model (LLM) serving poses major…
arXiv:2609.00588v3 Announce Type: replace Abstract: Reranking methods, such as Minimum Bayes Risk (MBR) decoding and Quality Estimation (QE) reranking…
arXiv:2608.16379v2 Announce Type: replace Abstract: Evaluating speech recognition for a Kurdish variety written in a Latin field orthography, using a…
arXiv:2608.11350v2 Announce Type: replace Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depe…
arXiv:2608.05906v2 Announce Type: replace Abstract: LLM agents that repair failures often discard successful corrections, forcing later episodes to re…
arXiv:2607.27232v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worl…
arXiv:2607.22376v2 Announce Type: replace Abstract: Most endangered languages lack the parallel data required for machine translation, despite the exi…
arXiv:2606.21844v2 Announce Type: replace Abstract: As AI systems integrate into online spaces, differentiating them from humans in conversations is i…
arXiv:2606.19468v2 Announce Type: replace Abstract: The narrative composition of web-scale LLM pretraining corpora remains largely unexplored, even th…
arXiv:2606.05134v2 Announce Type: replace Abstract: Deep active learning has previously been explored for LLM in-context sample selection, but not wit…
arXiv:2605.27865v2 Announce Type: replace Abstract: Matching submissions with suitable reviewers at scale is a growing challenge for major venues, yet…
arXiv:2605.21049v2 Announce Type: replace Abstract: Brain-language model alignment is often interpreted as evidence that transformer models implement…
arXiv:2605.18530v2 Announce Type: replace Abstract: While diffusion has drawn considerable recent attention from the language modeling community, cont…
arXiv:2605.12227v3 Announce Type: replace Abstract: Existing approaches to post-train models for long-context tasks face complementary limitations: (i…
arXiv:2605.00143v2 Announce Type: replace Abstract: Humor is a fundamental cognitive phenomenon in which humans derive pleasure from the expectation v…
arXiv:2604.25053v2 Announce Type: replace Abstract: While large language models (LLMs) are increasingly being explored for mental health applications,…
arXiv:2603.24125v3 Announce Type: replace Abstract: During training, Large Language Models (LLMs) learn social regularities that can lead to gender bi…
arXiv:2603.23821v2 Announce Type: replace Abstract: Linguistic representation learning in deep neural language models (LMs) has been studied for decad…