Back to AI Research

AI Research

GrammarRL trains models to produce useful answers within grammar constraints

Key Takeaways

  • Label-free reinforcement learning improves constrained greedy decoding on three structured tasks while moving search work into training.
  • A grammar can guarantee that a generated answer has an allowed structure while leaving its meaning wrong.
  • [GrammarRL](https://arxiv.org/abs/2609.39869), proposed by Gabriele Tuccio and colleagues, adapts language models to those constraints using reinforcement learning without annotated target outputs.
  • The method aims to improve the answer selected within the valid output space.
  • The authors evaluate sign-language gloss translation, hierarchical text classification and named-entity recognition with Llama models from one to eight billion parameters.

A grammar can guarantee that a generated answer has an allowed structure while leaving its meaning wrong. GrammarRL, proposed by Gabriele Tuccio and colleagues, adapts language models to those constraints using reinforcement learning without annotated target outputs. The method aims to improve the answer selected within the valid output space.
The authors evaluate sign-language gloss translation, hierarchical text classification and named-entity recognition with Llama models from one to eight billion parameters. They report an average improvement of 9.8 points over constrained greedy decoding, measured across the task-specific metrics. That average combines different metrics; it is not a universal accuracy percentage.

Valid tokens can still lead to an incorrect answer

Grammar-constrained decoding masks tokens that would violate the specified grammar. If the model assigns most probability to disallowed continuations, the remaining valid tokens can occupy a poorly learned part of its distribution. A locally preferred valid choice can then lead to a contextually wrong sequence.
Beam search explores several prefixes and can find better outputs, but its computation grows with beam width. High sequence likelihood is also an imperfect substitute for task quality. GrammarRL instead uses a tuning phase to shift the model toward useful valid sequences before deployment.
The grammar remains active when generating training candidates and at inference. Training changes the model's preferences among valid outputs; it does not discard the formal constraint. The experiments use structured JSON for classification and entity recognition, with closed vocabularies specific to those tasks.

Two likelihood rewards measure different properties

A frozen reference model scores each candidate in two directions. The direct reward measures the likelihood of the output given the input. The reverse reward measures how well the input can be reconstructed from that output, encouraging retention of source information.
Both rewards use unconstrained sequence probabilities. That avoids treating an unlikely sequence as strong evidence merely because masking has renormalized its few remaining token choices. The trainable policy adds LoRA adapters while preserving the base weights, and regularization limits departure from the reference model.
Training compares groups of sampled grammar-valid candidates, augmented with the top beam-search hypothesis, through a Reinforce Leave-One-Out objective. The beam candidate receives the same reward treatment as the sampled candidates; it is not an annotated answer that the policy must imitate.

Cheap inference still requires a tuning budget

The combined objective improves on constrained greedy decoding across all tested tasks and model sizes. The largest reported gain is 22.8 BLEU points on gloss translation. Single-reward versions sometimes beat the combined objective in an individual setting, but either reward alone can also underperform the frozen baseline.
The abstract reports matching or outperforming beam search on two of the three tasks while preserving greedy-decoding inference cost. That claim concerns inference, not the entire development budget: training still requires candidate generation, beam search and reference-model scoring.
A deployment should therefore evaluate semantic quality alongside grammar compliance and account for the cost of adaptation. These experiments support label-free tuning for the specified structured tasks. They do not establish that likelihood rewards guarantee factual correctness or that a grammar makes an unrestricted answer trustworthy.

Comments (0)

No comments yet

Be the first to share your thoughts!