When you need to serve a dozen different NLP tasks from a single endpoint, the usual answer is a massive multi‑task LLM like FLAN‑PaLM or OPT‑IML. Those models deliver impressive zero‑shot capabilities, but their size brings hefty hardware costs, memory constraints, and, often, closed‑source licences that block direct adaptation. What if you could keep a lean, open model in the loop and still reap the benefits of a giant‑scale LLM?
Enter Cappy, a 360 M‑parameter scorer trained on top of RoBERTa. Instead of generating text itself, Cappy evaluates an instruction–response pair and returns a confidence score between 0 and 1. The score can be used on its own for classification, or as a plug‑in that selects the best output from a larger language model. In this post we unpack how Cappy works, how it’s trained, and why it could become a go‑to component for developers building scalable, task‑aware AI services.
Why a Separate Scorer?
Traditional multi‑task LLMs follow an instruction‑following paradigm: each training example is turned into an instruction (e.g., “Combine the concepts ski, mountain, skier into a sentence”) and a target response (“Skier skis down the mountain”). The model learns to map the instruction to the response via teacher‑forcing, i.e., it is only ever exposed to the ground‑truth answer. This approach has two drawbacks for production use:
- Resource intensity – State‑of‑the‑art models range from 11 B to 540 B parameters, demanding multiple high‑end GPUs/TPUs for inference and making per‑task deployment expensive.
- Adaptation friction – Fine‑tuning or prompt‑tuning still requires back‑propagation through the LLM’s parameters, which keeps memory footprints high and prevents use with closed‑source APIs.
A lightweight scorer sidesteps both problems. It can be stored once, loaded on modest hardware, and applied to any LLM that can generate candidate responses. Because the scorer’s parameters stay fixed during downstream adaptation, developers avoid costly gradient passes on the large model.
Building Cappy: Data and Pre‑training
1. Gathering a Diverse Instruction Pool
Cappy’s foundation comes from the 39 datasets curated for PromptSource, the same collection used to train the T0 model. These datasets span question answering, sentiment analysis, summarization, and many other formats. Each example is converted into an instruction–ground‑truth pair via task‑specific templates.
2. Generating Weak Supervision Signals
For regression training, Cappy needs a continuous correctness label. The authors generated multiple candidate responses for each instruction by sampling from a strong multi‑task LLM (e.g., FLAN‑T5). Each candidate’s similarity to the ground truth was measured with Rouge‑L, yielding a score of 0–1 that serves as a weak supervision signal.
3. Massive Regression Corpus
The process produced 160 M instruction–response instances, each annotated with a Rouge‑L based correctness score. This corpus captures a spectrum of quality—from perfect matches to nonsensical outputs—enabling Cappy to learn a nuanced ranking function.
4. Continual Pre‑training on RoBERTa
Cappy is built on top of RoBERTa, a robust encoder‑only architecture. The regression dataset was used for continuous pre‑training on Google’s TPU‑v4 pods with the RedCoast framework, which streamlines distributed training. The final model retains only 360 M parameters, a fraction of the LLMs it will help evaluate.
How to Use Cappy in Practice
Candidate‑Selection Workflow
- Generate candidates – Use any language model (open‑source or API‑based) to produce n possible responses for a given instruction.
- Score each pair – Feed the instruction together with each candidate into Cappy. The model returns a probability‑like score.
- Pick the highest – Select the response with the largest score as the final answer.
For classification tasks, the set of possible answers (e.g., Yes/No for sentiment) is predefined, so Cappy can operate entirely on its own, without a generative model.
Adapting to Downstream Data
When you have task‑specific labeled data, you can fine‑tune Cappy on a small regression dataset built with the same Rouge‑L annotation pipeline. This step injects domain knowledge into the scorer while still leaving the backbone LLM frozen. Benefits include:
- Memory savings – No gradients flow through the large model.
- Closed‑source compatibility – Works with APIs that expose only inference capabilities.
- Unlimited supervision – Unlike in‑context learning, you can incorporate as many downstream examples as needed because scoring is performed per candidate, not concatenated to the prompt.
Empirical Results
Classification Benchmark (PromptSource)
Cappy (360 M) was evaluated on eleven held‑out tasks from PromptSource. It outperformed OPT‑175B and OPT‑IML‑30B, and its accuracy matched the best publicly available multi‑task LLMs such as T0‑11B and OPT‑IML‑175B. This demonstrates that a scorer can rival much larger models when used for answer selection.
Generation Benchmark (BIG‑Bench)
For the 45 generation‑only tasks in BIG‑Bench, FLAN‑T5 variants served as the base LLM. With Cappy scoring the sampled outputs, the average Rouge‑L score significantly increased across all model sizes, surpassing the strongest baseline that relied on self‑scoring via cross‑entropy. The improvement was especially pronounced for the largest FLAN‑T5‑XXL, confirming that Cappy scales well with stronger generators.
Key Takeaways for Developers
- Parameter efficiency – A 360 M scorer can replace or augment models that are hundreds of billions of parameters.
- Hardware‑friendly – Cappy runs comfortably on a single GPU/TPU core, cutting inference cost by orders of magnitude.
- Zero‑parameter access – Works with closed‑source models accessed via REST APIs, opening the door to commercial LLMs.
- Flexible supervision – You can fine‑tune the scorer on any amount of downstream data without touching the generator.
- Modular design – Cappy can be dropped into existing pipelines that already perform candidate generation, requiring only a lightweight scoring call.
Getting Started
- Install the pre‑trained scorer from the public model hub (e.g., via
transformers. - Wrap your generation code so that each output is paired with the original instruction.
- Call
cappy.score(instruction, candidate)and collect the confidence values. - Select the highest‑scoring candidate or, for classification, use the raw scores to make a decision.
- (Optional) Fine‑tune on a small regression set if you have task‑specific labels.
The entire workflow adds only a few milliseconds of latency per candidate, which is negligible compared to the cost of generating the candidates themselves.
Looking Ahead
Cappy showcases that a compact, pre‑trained scorer can both boost large language models and stand alone for certain tasks. Future research may explore:
- Extending the scorer to multi‑turn dialogues.
- Joint training of scorer and generator for end‑to‑end optimization.
- Using Cappy as a reward model in reinforcement learning‑from‑human‑feedback pipelines.
For developers who wrestle with the trade‑off between model performance and deployment cost, Cappy offers a pragmatic middle ground: high‑quality scoring without the heavyweight baggage of giant LLMs.
Source: Cappy: Outperforming and boosting large multi-task language models with a small scorer
Automated Transmission
This entry was synthesized and populated dynamically using native API integrations.