gemma-3-270m SFT'd on Skywork-Reward-Preference-80K (chosen-only)

A response-only SFT of google/gemma-3-270m on the chosen responses of Skywork/Skywork-Reward-Preference-80K-v0.2.

This is the anchor checkpoint used as the SFT reference policy for a research project on data-space DPO — applying preference signal by reweighting SFT samples through a susceptibility matrix, instead of via DPO's weight-space gradient descent. It is not intended as a useful chat model — at 270M parameters and a single epoch of SFT, generations are at best "coherent text on the topic of the prompt". The role of this checkpoint is to be a controlled, reproducible anchor for sampling-based experiments.

Training recipe

Base google/gemma-3-270m (pretrained, no chat template)
Data Skywork/Skywork-Reward-Preference-80K-v0.2, chosen responses only (~77K examples)
Epochs 1
Learning rate 2e-5, cosine schedule, 3% warmup
Batch size 16 per device × 4 GPUs = 64 effective
Precision bf16 + gradient checkpointing
Loss response-only cross-entropy (prompt tokens masked to -100)
Total steps 602
Final train loss 1.41
Wall time ~4 min on 8× H100

Prompt format

The base model has no chat_template, so we use a fixed plain-text format matching the SFT distribution:

Human: <user message>
Assistant: ASSISTANT_RESPONSE

Internally the prompt prefix is "Human: <user>\n\nAssistant:" and the loss is computed only on the tokens following Assistant:.

Quick start

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

mid = "bfpill/gemma-3-270m-sft-skywork-chosen"
tok = AutoTokenizer.from_pretrained(mid)
model = AutoModelForCausalLM.from_pretrained(mid, torch_dtype=torch.bfloat16).cuda().eval()

prompt = f"Human: What is the capital of France?\n\nAssistant:"
ids = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(
    **ids, max_new_tokens=80, do_sample=False,
    repetition_penalty=1.15, pad_token_id=tok.eos_token_id,
)
print(tok.decode(out[0, ids.input_ids.shape[1]:], skip_special_tokens=True))

Limitations

  • 270M params, 1 epoch of SFT on ~77K examples. Generations are short and often factually incorrect.
  • Trained exclusively on the Skywork prompt distribution; performance on other distributions (instruction-following, dialogue, code, math) reflects the base model plus a thin Skywork-flavoured veneer.
  • Not preference-tuned. The whole point of this checkpoint is to be the anchor for a follow-up data-space patterning experiment that is the replacement for the usual DPO step.
  • Inherits the base model's biases and license restrictions.

License

Gemma Terms of Use — see Google's Gemma license. This SFT preserves the base model's weights structure; usage is subject to the same terms.

Source code

The trainer that produced this checkpoint: rm-patterning/policy/sft_anchor.py in timaeus-research/timaeus@max/rm-scaling (commit 3a3e0755e and later).

torchrun --nproc_per_node=8 rm-patterning/policy/sft_anchor.py \
    --base-model google/gemma-3-270m \
    --output-dir ./gemma-3-270m-sft-skywork-chosen \
    --epochs 1 --learning-rate 2e-5 --batch-size 16 --max-length 1024
Downloads last month
17
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bfpill/gemma-3-270m-sft-skywork-chosen

Finetuned
(161)
this model

Dataset used to train bfpill/gemma-3-270m-sft-skywork-chosen