Summary
Add Customization Bench as a new evaluation in NeMo Gym. This benchmark measures how readily a base or GA model can be fine-tuned (SFT) to a specific domain (e.g., legal, finance) — assessing customizability as a first-class model property alongside raw capability.
Motivation
Frontier model comparisons today focus primarily on out-of-the-box performance. Customization Bench fills a gap by evaluating how efficiently a model adapts to domain-specific tasks with supervised fine-tuning, enabling apples-to-apples comparison of base models on their customizability.
Summary
Add Customization Bench as a new evaluation in NeMo Gym. This benchmark measures how readily a base or GA model can be fine-tuned (SFT) to a specific domain (e.g., legal, finance) — assessing customizability as a first-class model property alongside raw capability.
Motivation
Frontier model comparisons today focus primarily on out-of-the-box performance. Customization Bench fills a gap by evaluating how efficiently a model adapts to domain-specific tasks with supervised fine-tuning, enabling apples-to-apples comparison of base models on their customizability.