English

Model ReleasesHugging FaceLFM2.5-350M

Method for Improving Structured Output in 350M-Scale Models Using GRPO Released

This article is a translation. Read the Japanese original

Hugging Face has released a low-cost fine-tuning method to increase the accuracy of structured outputs for the LFM2.5-350M model, which has 350 million parameters.

This method utilizes the TRL library and Group Relative Policy Optimization (GRPO). The training requires approximately 500 samples and 100 steps, making it a scale that can be executed even on the free GPU tiers of Colab or Kaggle.

In evaluations using the IFStruct benchmark, it was reported that the base model's compliance rate of 22.6% improved to 29.7% after fine-tuning.

This initiative aims to enhance the model's ability to output in a parsable format according to the requested schema.


Source: Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps (Hugging Face Blog, 2026-09-03)