qwen2.5-1.5b-diffusion-chatmix-1024-2m-block-shift-v3

This model was trained from scratch on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 2.9872

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 3e-05
  • train_batch_size: 4
  • eval_batch_size: 2
  • seed: 42
  • gradient_accumulation_steps: 32
  • total_train_batch_size: 128
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 500.0
  • num_epochs: 1.0

Training results

Training Loss Epoch Step Validation Loss
4.4160 0.0149 200 4.4659
3.8561 0.0297 400 3.9072
3.7444 0.0446 600 3.6386
3.5010 0.0594 800 3.4406
3.1976 0.0743 1000 3.5046
3.2365 0.0891 1200 3.2334
3.4797 0.1040 1400 3.3359
3.0409 0.1188 1600 3.1277
3.1039 0.1337 1800 2.9810
2.9802 0.1485 2000 3.0258
2.9370 0.1634 2200 3.1481
2.8930 0.1782 2400 3.0833
3.0009 0.1931 2600 2.9747
3.0520 0.2079 2800 3.0349
2.9531 0.2228 3000 2.8790
3.0516 0.2376 3200 2.9013
2.8841 0.2525 3400 2.9984
2.7605 0.2673 3600 3.0805
2.7315 0.2822 3800 2.8539
2.8823 0.2970 4000 3.0194
2.8969 0.3119 4200 2.9479
2.8563 0.3267 4400 2.7973
2.7066 0.3416 4600 3.0403
2.9566 0.3565 4800 2.7825
2.7957 0.3713 5000 2.7929
2.7575 0.3862 5200 2.8629
2.8845 0.4010 5400 2.8028
2.7579 0.4159 5600 2.8800
2.7875 0.4307 5800 2.7144
2.7243 0.4456 6000 2.7281
2.7424 0.4604 6200 2.8936
2.7749 0.4753 6400 2.8784
2.7307 0.4901 6600 2.9189
2.7434 0.5050 6800 2.9361
2.8430 0.5198 7000 2.8152
2.8007 0.5347 7200 2.7186
2.7972 0.5495 7400 2.6337
2.7435 0.5644 7600 2.7856
2.6787 0.5792 7800 2.7218
2.7944 0.5941 8000 2.8605
2.8321 0.6089 8200 2.9186
2.8288 0.6238 8400 2.6986
2.7419 0.6386 8600 2.9237
2.6824 0.6535 8800 2.9058
2.5941 0.6683 9000 2.7429
2.6812 0.6832 9200 2.7436
2.8702 0.6981 9400 2.8215
2.6919 0.7129 9600 2.6216
2.6482 0.7278 9800 2.9565
2.6567 0.7426 10000 2.7712
2.5969 0.7575 10200 2.8750
2.6823 0.7723 10400 3.0409
2.7710 0.7872 10600 2.6757
2.6226 0.8020 10800 2.6895
2.8287 0.8169 11000 2.8124
2.7751 0.8317 11200 2.7178
2.7334 0.8466 11400 2.7980
2.8281 0.8614 11600 2.6867
2.7158 0.8763 11800 2.7768
2.6976 0.8911 12000 2.6079
2.8079 0.9060 12200 2.6920
2.7473 0.9208 12400 2.9175
2.7127 0.9357 12600 2.7127
2.7303 0.9505 12800 2.9447
2.5962 0.9654 13000 2.6765
2.6588 0.9802 13200 2.8538
2.6086 0.9951 13400 2.5556
2.7962 1.0 13466 2.9872

Framework versions

  • Transformers 5.14.1
  • Pytorch 2.12.1+cu130
  • Datasets 4.8.5
  • Tokenizers 0.22.2
Downloads last month
24
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support