A model can choose the right legal answer from four options and still struggle to write a dependable explanation. Rubén Manrique and colleagues examine that gap in a benchmark of Colombian law, where language, legal terminology and the applicable legal order require jurisdiction-specific evaluation.
Their Colombian Legal Reliability Benchmark study evaluates 15 models on 1,042 expert-validated items across ten areas of law. The paper reports high multiple-choice accuracy for some models alongside substantially weaker factual correctness in free-text responses. These are separate metrics on different question subsets, not interchangeable percentages of legal advice that can be trusted.
Recognition and generation are separate tests
The benchmark contains 305 closed multiple-choice questions, 682 semi-open questions and 55 open-ended questions. Semi-open items request concise legal answers at different complexity levels. Open-ended cases have expert reference answers organized around issue, rule, application and conclusion.
The ten areas include constitutional, criminal and administrative law, alongside subjects such as family and tax law. Contributors select subjects matching their expertise. Every item passes the same review route regardless of whether students, workshops or the legal team supplied it.
The project receives 1,084 questions and archives 42 during validation. Of the accepted questions, 93% undergo at least one edit before approval. That supports the description of active curation, while the paper also acknowledges that the protocol lacks independent double coding with a formal inter-annotator agreement statistic.
Fluent responses are not a correctness measure
Closed-question accuracy ranges from 0.905 to 0.577 across the tested models. The top closed score belongs to Gemini 3.1 Pro. Free-text factual correctness does not exceed 0.45 for any model on the study's zero-to-one scoring scale.
The authors report a negative rank correlation of -0.46 between answer relevancy and correctness. Responses can address the question in a convincing style while conflicting with the expert reference. Multiple-choice accuracy and free-text correctness have a much stronger positive rank correlation of 0.94, suggesting that cheap closed-question screening can help compare model ordering while overstating absolute reliability.
Free-text evaluation combines reference-based metrics with separate judge and human validation procedures. The paper reports that an independent rubric-based judge and blind expert scoring reproduce the free-text ranking. Its abstract also reports that only about half of the cited norms are correct, with the remainder wrong or nonexistent. That is a reported benchmark finding, not an audit Franklin conducted of legal citations.
The evaluation excludes live retrieval
The tested models answer without internet access or external tools. Their scores describe responses based on internalized knowledge under the paper's prompts, not a lawyer using an assistant connected to an authoritative legal database.
Grounding in current sources is a proposed route to improvement, but this experiment does not establish how much a particular retrieval product would improve. Legal area and question complexity also affect results, making a single overall model score an incomplete purchasing or practice decision.
The paper gives Colombian legal practitioners a more relevant evaluation than borrowing a U.S. benchmark. Its results support expert supervision and citation verification. They do not support replacing professional judgment with the model that leads one multiple-choice table.
Comments