Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System

Administrator 0 阅读

AI Digest - ArXiv AI

Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System

Large language models (LLMs) are increasingly used to support legal practice, education, and research, yet their reliability in national legal systems outside the United States remains largely undocumented. We introduce an expert-validated benchmark for evaluating LLM reliability on the Colombian legal system. The benchmark comprises 1,042 items spanning ten areas of law and three question formats (closed multiple-choice, semi-open, and open-ended IRAC), built through a human-in-the-loop pipelin


Source: ArXiv AI