Latest AI/ML News

2784 articles · arXiv cs.CL

arXiv cs.CLOct 7, 2026

arXiv:2610.08747v1 Announce Type: new Abstract: A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two…

arXiv cs.CLOct 7, 2026

arXiv:2610.08719v1 Announce Type: new Abstract: Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward. We study evalua…

arXiv cs.CLOct 7, 2026

arXiv:2610.08703v1 Announce Type: new Abstract: In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners' problem-solvi…

arXiv cs.CLOct 7, 2026

arXiv:2610.08675v1 Announce Type: new Abstract: Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generate…

arXiv cs.CLOct 7, 2026

arXiv:2610.08660v1 Announce Type: new Abstract: Background: Biomedical AI can generate plausible explanations without reliably verifying whether each…

arXiv cs.CLOct 7, 2026

arXiv:2610.08630v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge…

arXiv cs.CLOct 7, 2026

arXiv:2610.08604v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, an…

arXiv cs.CLOct 7, 2026

arXiv:2610.08601v1 Announce Type: new Abstract: Emergency messaging such as extreme-weather reports and earthquake instructions can involve high stake…

arXiv cs.CLOct 7, 2026

arXiv:2610.08585v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinica…

arXiv cs.CLOct 7, 2026

arXiv:2610.08544v1 Announce Type: new Abstract: Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model…

arXiv cs.CLOct 7, 2026

arXiv:2610.08513v1 Announce Type: new Abstract: LLMs are increasingly deployed as autonomous agents in social environments, making it critical to stud…

arXiv cs.CLOct 7, 2026

arXiv:2610.08501v1 Announce Type: new Abstract: Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but…

arXiv cs.CLOct 7, 2026

arXiv:2610.08448v1 Announce Type: new Abstract: On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With diff…

arXiv cs.CLOct 7, 2026

arXiv:2610.08388v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they sti…

arXiv cs.CLOct 7, 2026

arXiv:2610.08303v1 Announce Type: new Abstract: Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomo…

arXiv cs.CLOct 7, 2026

arXiv:2610.08300v1 Announce Type: new Abstract: Long-term conversational memory is becoming an integral component of modern LLM systems. Proposed arch…

arXiv cs.CLOct 7, 2026

arXiv:2610.08208v1 Announce Type: new Abstract: We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations…

arXiv cs.CLOct 7, 2026

arXiv:2610.08153v1 Announce Type: new Abstract: Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the…

arXiv cs.CLOct 7, 2026

arXiv:2610.08093v1 Announce Type: new Abstract: Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely co…

arXiv cs.CLOct 7, 2026

arXiv:2610.08085v1 Announce Type: new Abstract: Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognit…