A banking presentation can contain fluent recommendations while its figures disagree across slides. Ludovic Gibert, Matis Despujols and André-Louis Rochet study a generation workflow that combines a 27B language model with financial calculations, templates and validation checks. Their banking presentation preprint reports substantial gains over the same model answering a short prompt, alongside a measurement problem: AI judges change their scores even when the decks stay identical.
The authors retrospectively examine seventeen development deliverables spanning credit reviews, financing pitches and transaction advice. These are exercises built around listed or rated issuers and public information. The evaluation begins with collected inputs; it does not measure the wider application's information collection or banker approval process.
Calculations and templates carry part of the workload
The system uses an on-premises Qwen3.8-27B model alongside a typed registry of financial facts. Entries include values, units, periods and provenance, with sources or formulas where applicable. Missing bank-held information becomes a named gap. Playbooks determine the page plan, and model roles contribute recommendations or narrative choices.
Some pages can use compiled prose without a writer call. After generation, validation rejects unsupported numeric literals and checks contradictions and layout defects. This is an output check rather than constrained decoding. Historical archives do not connect every evaluated deck to a complete code, input and generation-log snapshot, so the authors cannot quantify how much of its prose the model wrote.
That separation matters when interpreting an improvement. A higher grade belongs to the combined system, including its calculations and templates. The experiment cannot assign the whole gain to the language model or establish which individual component caused it.
Large gains survive a text-only comparison
In shared-session text-only grading with template markers removed, five judges scored the complete system 20.4 to 33.6 points out of 95 above direct generation by the same model from a short prompt. Every judge preferred the system on all seventeen deliverables in that comparison.
Against direct Opus generation from a short prompt, margins ranged from minus 4.7 to plus 0.8 points. Those results do not establish equivalence. The comparison also pits an iteratively developed workflow against single-sample direct generations on its development cases, rather than demonstrating generalization to reserved banking tasks.
The authors used a broader panel of eight judges from six model families. Judges agreed on progress across development rounds more than on the ranking of final decks. Longer system outputs leave length-related grading bias possible, even after removing slides and template identifiers.
Repeated grades need a reproducible judge configuration
Two Opus passes on unchanged decks shifted average scores by 2.0 points in one round and 1.1 in another. Across those repeated assessments, the estimated per-deck noise scale was about 7.9 points under the study's statistical assumptions. Decks graded in the same session appeared to share offsets, weakening the case for treating every score as an independent measurement.
The study has no banker-rated reference set. Judge agreement therefore cannot establish professional acceptability or financial accuracy. Teams using AI feedback to decide whether a deck improved should preserve the grading configuration, repeat assessments and distinguish a large workflow advantage from a small score movement within grading variability.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!