arXiv cs.AIOctober 7, 2026
MERCI Cards: An LLM Evaluation and Deployment Framework for High-Stakes Domains
Excerpt
arXiv:2610.04430v1 Announce Type: new Abstract: As LLMs are increasingly deployed in high-stakes professional workflows, engineers and researchers require principled protocols to systematically track, monitor, and improve model performance across deployment cycles. We present a mathematical framework for iterative LLM evaluation and deployment, and demonstrate its application to AI systems used in criminal justice. Our framework formalizes LLM integration in high-stakes, high-risk, and resource-