adugeen commited on
Commit
7615637
·
verified ·
1 Parent(s): 1fddebe

Upload folder using huggingface_hub

Browse files
1_Pooling/config.json CHANGED
@@ -1,5 +1,5 @@
1
  {
2
- "word_embedding_dimension": 384,
3
  "pooling_mode_cls_token": false,
4
  "pooling_mode_mean_tokens": true,
5
  "pooling_mode_max_tokens": false,
 
1
  {
2
+ "word_embedding_dimension": 768,
3
  "pooling_mode_cls_token": false,
4
  "pooling_mode_mean_tokens": true,
5
  "pooling_mode_max_tokens": false,
README.md CHANGED
@@ -6,7 +6,7 @@ tags:
6
  - generated_from_trainer
7
  - dataset_size:276686
8
  - loss:MultipleNegativesRankingLoss
9
- base_model: intfloat/multilingual-e5-small
10
  widget:
11
  - source_sentence: 'query: Печенеги отступали. Они могли запросто убить оставшегося
12
  позади князя Владимира, мальчишку и старика, но получили приказ - уходить. Куря
@@ -2907,7 +2907,7 @@ metrics:
2907
  - cosine_ap
2908
  - cosine_mcc
2909
  model-index:
2910
- - name: SentenceTransformer based on intfloat/multilingual-e5-small
2911
  results:
2912
  - task:
2913
  type: binary-classification
@@ -2917,42 +2917,42 @@ model-index:
2917
  type: unknown
2918
  metrics:
2919
  - type: cosine_accuracy
2920
- value: 0.9208363155269265
2921
  name: Cosine Accuracy
2922
  - type: cosine_accuracy_threshold
2923
- value: 0.8324185609817505
2924
  name: Cosine Accuracy Threshold
2925
  - type: cosine_f1
2926
- value: 0.7493691493691494
2927
  name: Cosine F1
2928
  - type: cosine_f1_threshold
2929
- value: 0.8252550363540649
2930
  name: Cosine F1 Threshold
2931
  - type: cosine_precision
2932
- value: 0.7499918532277512
2933
  name: Cosine Precision
2934
  - type: cosine_recall
2935
- value: 0.7487474786908712
2936
  name: Cosine Recall
2937
  - type: cosine_ap
2938
- value: 0.8415735890036254
2939
  name: Cosine Ap
2940
  - type: cosine_mcc
2941
- value: 0.6992932604835177
2942
  name: Cosine Mcc
2943
  ---
2944
 
2945
- # SentenceTransformer based on intfloat/multilingual-e5-small
2946
 
2947
- This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [intfloat/multilingual-e5-small](https://huggingface.co/intfloat/multilingual-e5-small) on the json dataset. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
2948
 
2949
  ## Model Details
2950
 
2951
  ### Model Description
2952
  - **Model Type:** Sentence Transformer
2953
- - **Base model:** [intfloat/multilingual-e5-small](https://huggingface.co/intfloat/multilingual-e5-small) <!-- at revision c007d7ef6fd86656326059b28395a7a03a7c5846 -->
2954
  - **Maximum Sequence Length:** 512 tokens
2955
- - **Output Dimensionality:** 384 dimensions
2956
  - **Similarity Function:** Cosine Similarity
2957
  - **Training Dataset:**
2958
  - json
@@ -2969,8 +2969,8 @@ This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [i
2969
 
2970
  ```
2971
  SentenceTransformer(
2972
- (0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: BertModel
2973
- (1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
2974
  (2): Normalize()
2975
  )
2976
  ```
@@ -2999,7 +2999,7 @@ sentences = [
2999
  ]
3000
  embeddings = model.encode(sentences)
3001
  print(embeddings.shape)
3002
- # [3, 384]
3003
 
3004
  # Get the similarity scores for the embeddings
3005
  similarities = model.similarity(embeddings, embeddings)
@@ -3041,14 +3041,14 @@ You can finetune this model on your own dataset.
3041
 
3042
  | Metric | Value |
3043
  |:--------------------------|:-----------|
3044
- | cosine_accuracy | 0.9208 |
3045
- | cosine_accuracy_threshold | 0.8324 |
3046
- | cosine_f1 | 0.7494 |
3047
- | cosine_f1_threshold | 0.8253 |
3048
- | cosine_precision | 0.75 |
3049
- | cosine_recall | 0.7487 |
3050
- | **cosine_ap** | **0.8416** |
3051
- | cosine_mcc | 0.6993 |
3052
 
3053
  <!--
3054
  ## Bias, Risks and Limitations
@@ -3120,8 +3120,8 @@ You can finetune this model on your own dataset.
3120
  #### Non-Default Hyperparameters
3121
 
3122
  - `eval_strategy`: epoch
3123
- - `per_device_train_batch_size`: 136
3124
- - `per_device_eval_batch_size`: 136
3125
  - `weight_decay`: 0.01
3126
  - `num_train_epochs`: 5
3127
  - `bf16`: True
@@ -3135,8 +3135,8 @@ You can finetune this model on your own dataset.
3135
  - `do_predict`: False
3136
  - `eval_strategy`: epoch
3137
  - `prediction_loss_only`: True
3138
- - `per_device_train_batch_size`: 136
3139
- - `per_device_eval_batch_size`: 136
3140
  - `per_gpu_train_batch_size`: None
3141
  - `per_gpu_eval_batch_size`: None
3142
  - `gradient_accumulation_steps`: 1
@@ -3248,72 +3248,96 @@ You can finetune this model on your own dataset.
3248
  </details>
3249
 
3250
  ### Training Logs
3251
- | Epoch | Step | Training Loss | Validation Loss | cosine_ap |
3252
- |:------:|:----:|:-------------:|:---------------:|:---------:|
3253
- | 0.0491 | 100 | 3.1392 | - | - |
3254
- | 0.0983 | 200 | 2.8632 | - | - |
3255
- | 0.1474 | 300 | 2.6963 | - | - |
3256
- | 0.1966 | 400 | 2.6594 | - | - |
3257
- | 0.2457 | 500 | 2.584 | - | - |
3258
- | 0.2948 | 600 | 2.5429 | - | - |
3259
- | 0.3440 | 700 | 2.5105 | - | - |
3260
- | 0.3931 | 800 | 2.4482 | - | - |
3261
- | 0.4423 | 900 | 2.442 | - | - |
3262
- | 0.4914 | 1000 | 2.4386 | - | - |
3263
- | 0.5405 | 1100 | 2.3912 | - | - |
3264
- | 0.5897 | 1200 | 2.3778 | - | - |
3265
- | 0.6388 | 1300 | 2.3481 | - | - |
3266
- | 0.6880 | 1400 | 2.3325 | - | - |
3267
- | 0.7371 | 1500 | 2.3082 | - | - |
3268
- | 0.7862 | 1600 | 2.2942 | - | - |
3269
- | 0.8354 | 1700 | 2.3115 | - | - |
3270
- | 0.8845 | 1800 | 2.2769 | - | - |
3271
- | 0.9337 | 1900 | 2.2402 | - | - |
3272
- | 0.9828 | 2000 | 2.2372 | - | - |
3273
- | 1.0 | 2035 | - | 8.7903 | 0.8367 |
3274
- | 1.0319 | 2100 | 2.1558 | - | - |
3275
- | 1.0811 | 2200 | 2.0741 | - | - |
3276
- | 1.1302 | 2300 | 2.0783 | - | - |
3277
- | 1.1794 | 2400 | 2.0596 | - | - |
3278
- | 1.2285 | 2500 | 2.0398 | - | - |
3279
- | 1.2776 | 2600 | 2.0504 | - | - |
3280
- | 1.3268 | 2700 | 2.0674 | - | - |
3281
- | 1.3759 | 2800 | 2.0305 | - | - |
3282
- | 1.4251 | 2900 | 2.0372 | - | - |
3283
- | 1.4742 | 3000 | 2.0271 | - | - |
3284
- | 1.5233 | 3100 | 2.007 | - | - |
3285
- | 1.5725 | 3200 | 2.0074 | - | - |
3286
- | 1.6216 | 3300 | 1.9978 | - | - |
3287
- | 1.6708 | 3400 | 1.974 | - | - |
3288
- | 1.7199 | 3500 | 1.9922 | - | - |
3289
- | 1.7690 | 3600 | 1.9743 | - | - |
3290
- | 1.8182 | 3700 | 1.9536 | - | - |
3291
- | 1.8673 | 3800 | 1.9717 | - | - |
3292
- | 1.9165 | 3900 | 1.9324 | - | - |
3293
- | 1.9656 | 4000 | 1.9275 | - | - |
3294
- | 2.0 | 4070 | - | 9.2237 | 0.8386 |
3295
- | 2.0147 | 4100 | 1.9059 | - | - |
3296
- | 2.0639 | 4200 | 1.7814 | - | - |
3297
- | 2.1130 | 4300 | 1.7528 | - | - |
3298
- | 2.1622 | 4400 | 1.786 | - | - |
3299
- | 2.2113 | 4500 | 1.7963 | - | - |
3300
- | 2.2604 | 4600 | 1.7744 | - | - |
3301
- | 2.3096 | 4700 | 1.7753 | - | - |
3302
- | 2.3587 | 4800 | 1.7671 | - | - |
3303
- | 2.4079 | 4900 | 1.7832 | - | - |
3304
- | 2.4570 | 5000 | 1.7715 | - | - |
3305
- | 2.5061 | 5100 | 1.721 | - | - |
3306
- | 2.5553 | 5200 | 1.7584 | - | - |
3307
- | 2.6044 | 5300 | 1.7348 | - | - |
3308
- | 2.6536 | 5400 | 1.7331 | - | - |
3309
- | 2.7027 | 5500 | 1.7274 | - | - |
3310
- | 2.7518 | 5600 | 1.7587 | - | - |
3311
- | 2.8010 | 5700 | 1.7379 | - | - |
3312
- | 2.8501 | 5800 | 1.7579 | - | - |
3313
- | 2.8993 | 5900 | 1.7573 | - | - |
3314
- | 2.9484 | 6000 | 1.708 | - | - |
3315
- | 2.9975 | 6100 | 1.7159 | - | - |
3316
- | 3.0 | 6105 | - | 9.8395 | 0.8416 |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3317
 
3318
 
3319
  ### Framework Versions
 
6
  - generated_from_trainer
7
  - dataset_size:276686
8
  - loss:MultipleNegativesRankingLoss
9
+ base_model: intfloat/multilingual-e5-base
10
  widget:
11
  - source_sentence: 'query: Печенеги отступали. Они могли запросто убить оставшегося
12
  позади князя Владимира, мальчишку и старика, но получили приказ - уходить. Куря
 
2907
  - cosine_ap
2908
  - cosine_mcc
2909
  model-index:
2910
+ - name: SentenceTransformer based on intfloat/multilingual-e5-base
2911
  results:
2912
  - task:
2913
  type: binary-classification
 
2917
  type: unknown
2918
  metrics:
2919
  - type: cosine_accuracy
2920
+ value: 0.9225280326197758
2921
  name: Cosine Accuracy
2922
  - type: cosine_accuracy_threshold
2923
+ value: 0.7901061773300171
2924
  name: Cosine Accuracy Threshold
2925
  - type: cosine_f1
2926
+ value: 0.7559554803436604
2927
  name: Cosine F1
2928
  - type: cosine_f1_threshold
2929
+ value: 0.7817596793174744
2930
  name: Cosine F1 Threshold
2931
  - type: cosine_precision
2932
+ value: 0.756201575623413
2933
  name: Cosine Precision
2934
  - type: cosine_recall
2935
+ value: 0.7557095451883662
2936
  name: Cosine Recall
2937
  - type: cosine_ap
2938
+ value: 0.8478615501518483
2939
  name: Cosine Ap
2940
  - type: cosine_mcc
2941
+ value: 0.7071656901034916
2942
  name: Cosine Mcc
2943
  ---
2944
 
2945
+ # SentenceTransformer based on intfloat/multilingual-e5-base
2946
 
2947
+ This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [intfloat/multilingual-e5-base](https://huggingface.co/intfloat/multilingual-e5-base) on the json dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
2948
 
2949
  ## Model Details
2950
 
2951
  ### Model Description
2952
  - **Model Type:** Sentence Transformer
2953
+ - **Base model:** [intfloat/multilingual-e5-base](https://huggingface.co/intfloat/multilingual-e5-base) <!-- at revision 835193815a3936a24a0ee7dc9e3d48c1fbb19c55 -->
2954
  - **Maximum Sequence Length:** 512 tokens
2955
+ - **Output Dimensionality:** 768 dimensions
2956
  - **Similarity Function:** Cosine Similarity
2957
  - **Training Dataset:**
2958
  - json
 
2969
 
2970
  ```
2971
  SentenceTransformer(
2972
+ (0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: XLMRobertaModel
2973
+ (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
2974
  (2): Normalize()
2975
  )
2976
  ```
 
2999
  ]
3000
  embeddings = model.encode(sentences)
3001
  print(embeddings.shape)
3002
+ # [3, 768]
3003
 
3004
  # Get the similarity scores for the embeddings
3005
  similarities = model.similarity(embeddings, embeddings)
 
3041
 
3042
  | Metric | Value |
3043
  |:--------------------------|:-----------|
3044
+ | cosine_accuracy | 0.9225 |
3045
+ | cosine_accuracy_threshold | 0.7901 |
3046
+ | cosine_f1 | 0.756 |
3047
+ | cosine_f1_threshold | 0.7818 |
3048
+ | cosine_precision | 0.7562 |
3049
+ | cosine_recall | 0.7557 |
3050
+ | **cosine_ap** | **0.8479** |
3051
+ | cosine_mcc | 0.7072 |
3052
 
3053
  <!--
3054
  ## Bias, Risks and Limitations
 
3120
  #### Non-Default Hyperparameters
3121
 
3122
  - `eval_strategy`: epoch
3123
+ - `per_device_train_batch_size`: 64
3124
+ - `per_device_eval_batch_size`: 64
3125
  - `weight_decay`: 0.01
3126
  - `num_train_epochs`: 5
3127
  - `bf16`: True
 
3135
  - `do_predict`: False
3136
  - `eval_strategy`: epoch
3137
  - `prediction_loss_only`: True
3138
+ - `per_device_train_batch_size`: 64
3139
+ - `per_device_eval_batch_size`: 64
3140
  - `per_gpu_train_batch_size`: None
3141
  - `per_gpu_eval_batch_size`: None
3142
  - `gradient_accumulation_steps`: 1
 
3248
  </details>
3249
 
3250
  ### Training Logs
3251
+ | Epoch | Step | Training Loss | Validation Loss | cosine_ap |
3252
+ |:------:|:-----:|:-------------:|:---------------:|:---------:|
3253
+ | 1.0176 | 4400 | 1.4186 | - | - |
3254
+ | 1.0407 | 4500 | 1.4075 | - | - |
3255
+ | 1.0638 | 4600 | 1.3934 | - | - |
3256
+ | 1.0870 | 4700 | 1.3799 | - | - |
3257
+ | 1.1101 | 4800 | 1.3597 | - | - |
3258
+ | 1.1332 | 4900 | 1.3351 | - | - |
3259
+ | 1.1563 | 5000 | 1.3082 | - | - |
3260
+ | 1.1795 | 5100 | 1.3105 | - | - |
3261
+ | 1.2026 | 5200 | 1.2948 | - | - |
3262
+ | 1.2257 | 5300 | 1.3486 | - | - |
3263
+ | 1.2488 | 5400 | 1.3155 | - | - |
3264
+ | 1.2720 | 5500 | 1.2761 | - | - |
3265
+ | 1.2951 | 5600 | 1.2541 | - | - |
3266
+ | 1.3182 | 5700 | 1.2346 | - | - |
3267
+ | 1.3414 | 5800 | 1.2285 | - | - |
3268
+ | 1.3645 | 5900 | 1.2013 | - | - |
3269
+ | 1.3876 | 6000 | 1.1986 | - | - |
3270
+ | 1.4107 | 6100 | 1.1755 | - | - |
3271
+ | 1.4339 | 6200 | 1.1937 | - | - |
3272
+ | 1.4570 | 6300 | 1.202 | - | - |
3273
+ | 1.4801 | 6400 | 1.1607 | - | - |
3274
+ | 1.5032 | 6500 | 1.2116 | - | - |
3275
+ | 1.5264 | 6600 | 1.1797 | - | - |
3276
+ | 1.5495 | 6700 | 1.1571 | - | - |
3277
+ | 1.5726 | 6800 | 1.1526 | - | - |
3278
+ | 1.5957 | 6900 | 1.1438 | - | - |
3279
+ | 1.6189 | 7000 | 1.1634 | - | - |
3280
+ | 1.6420 | 7100 | 1.1367 | - | - |
3281
+ | 1.6651 | 7200 | 1.1133 | - | - |
3282
+ | 1.6883 | 7300 | 1.1156 | - | - |
3283
+ | 1.7114 | 7400 | 1.1102 | - | - |
3284
+ | 1.7345 | 7500 | 1.1123 | - | - |
3285
+ | 1.7576 | 7600 | 1.1066 | - | - |
3286
+ | 1.7808 | 7700 | 1.1291 | - | - |
3287
+ | 1.8039 | 7800 | 1.1094 | - | - |
3288
+ | 1.8270 | 7900 | 1.094 | - | - |
3289
+ | 1.8501 | 8000 | 1.1585 | - | - |
3290
+ | 1.8733 | 8100 | 1.077 | - | - |
3291
+ | 1.8964 | 8200 | 1.108 | - | - |
3292
+ | 1.9195 | 8300 | 1.1431 | - | - |
3293
+ | 1.9426 | 8400 | 1.0784 | - | - |
3294
+ | 1.9658 | 8500 | 1.0834 | - | - |
3295
+ | 1.9889 | 8600 | 1.1268 | - | - |
3296
+ | 2.0 | 8648 | - | 9.6992 | 0.8450 |
3297
+ | 2.0120 | 8700 | 1.0443 | - | - |
3298
+ | 2.0352 | 8800 | 0.9715 | - | - |
3299
+ | 2.0583 | 8900 | 0.957 | - | - |
3300
+ | 2.0814 | 9000 | 0.9784 | - | - |
3301
+ | 2.1045 | 9100 | 0.9581 | - | - |
3302
+ | 2.1277 | 9200 | 0.9569 | - | - |
3303
+ | 2.1508 | 9300 | 0.9518 | - | - |
3304
+ | 2.1739 | 9400 | 0.9485 | - | - |
3305
+ | 2.1970 | 9500 | 0.9433 | - | - |
3306
+ | 2.2202 | 9600 | 0.9392 | - | - |
3307
+ | 2.2433 | 9700 | 0.9248 | - | - |
3308
+ | 2.2664 | 9800 | 0.9105 | - | - |
3309
+ | 2.2895 | 9900 | 0.9769 | - | - |
3310
+ | 2.3127 | 10000 | 0.9502 | - | - |
3311
+ | 2.3358 | 10100 | 0.9604 | - | - |
3312
+ | 2.3589 | 10200 | 0.9291 | - | - |
3313
+ | 2.3821 | 10300 | 0.9552 | - | - |
3314
+ | 2.4052 | 10400 | 0.9621 | - | - |
3315
+ | 2.4283 | 10500 | 0.9357 | - | - |
3316
+ | 2.4514 | 10600 | 0.9323 | - | - |
3317
+ | 2.4746 | 10700 | 0.9327 | - | - |
3318
+ | 2.4977 | 10800 | 0.9067 | - | - |
3319
+ | 2.5208 | 10900 | 0.9411 | - | - |
3320
+ | 2.5439 | 11000 | 0.9305 | - | - |
3321
+ | 2.5671 | 11100 | 0.9378 | - | - |
3322
+ | 2.5902 | 11200 | 0.9171 | - | - |
3323
+ | 2.6133 | 11300 | 0.9074 | - | - |
3324
+ | 2.6364 | 11400 | 0.9262 | - | - |
3325
+ | 2.6596 | 11500 | 0.9063 | - | - |
3326
+ | 2.6827 | 11600 | 0.8814 | - | - |
3327
+ | 2.7058 | 11700 | 0.9089 | - | - |
3328
+ | 2.7290 | 11800 | 0.9048 | - | - |
3329
+ | 2.7521 | 11900 | 0.9268 | - | - |
3330
+ | 2.7752 | 12000 | 0.8913 | - | - |
3331
+ | 2.7983 | 12100 | 0.9064 | - | - |
3332
+ | 2.8215 | 12200 | 0.8585 | - | - |
3333
+ | 2.8446 | 12300 | 0.878 | - | - |
3334
+ | 2.8677 | 12400 | 0.8612 | - | - |
3335
+ | 2.8908 | 12500 | 0.8799 | - | - |
3336
+ | 2.9140 | 12600 | 0.8541 | - | - |
3337
+ | 2.9371 | 12700 | 0.8521 | - | - |
3338
+ | 2.9602 | 12800 | 0.8582 | - | - |
3339
+ | 2.9833 | 12900 | 0.869 | - | - |
3340
+ | 3.0 | 12972 | - | 10.4115 | 0.8479 |
3341
 
3342
 
3343
  ### Framework Versions
config.json CHANGED
@@ -1,25 +1,27 @@
1
  {
2
  "architectures": [
3
- "BertModel"
4
  ],
5
  "attention_probs_dropout_prob": 0.1,
 
6
  "classifier_dropout": null,
 
7
  "hidden_act": "gelu",
8
  "hidden_dropout_prob": 0.1,
9
- "hidden_size": 384,
10
  "initializer_range": 0.02,
11
- "intermediate_size": 1536,
12
- "layer_norm_eps": 1e-12,
13
- "max_position_embeddings": 512,
14
- "model_type": "bert",
15
  "num_attention_heads": 12,
16
  "num_hidden_layers": 12,
17
- "pad_token_id": 0,
 
18
  "position_embedding_type": "absolute",
19
- "tokenizer_class": "XLMRobertaTokenizer",
20
  "torch_dtype": "float32",
21
  "transformers_version": "4.52.4",
22
- "type_vocab_size": 2,
23
  "use_cache": true,
24
- "vocab_size": 250037
25
  }
 
1
  {
2
  "architectures": [
3
+ "XLMRobertaModel"
4
  ],
5
  "attention_probs_dropout_prob": 0.1,
6
+ "bos_token_id": 0,
7
  "classifier_dropout": null,
8
+ "eos_token_id": 2,
9
  "hidden_act": "gelu",
10
  "hidden_dropout_prob": 0.1,
11
+ "hidden_size": 768,
12
  "initializer_range": 0.02,
13
+ "intermediate_size": 3072,
14
+ "layer_norm_eps": 1e-05,
15
+ "max_position_embeddings": 514,
16
+ "model_type": "xlm-roberta",
17
  "num_attention_heads": 12,
18
  "num_hidden_layers": 12,
19
+ "output_past": true,
20
+ "pad_token_id": 1,
21
  "position_embedding_type": "absolute",
 
22
  "torch_dtype": "float32",
23
  "transformers_version": "4.52.4",
24
+ "type_vocab_size": 1,
25
  "use_cache": true,
26
+ "vocab_size": 250002
27
  }
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:11407a208a2be3cd06b48f6a28931de116570bfb7520f75ae45cd99aa60d13c1
3
- size 470637416
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7224289aa386ecff5a6a47a5878ba9a3a59c124abf579c211a99dcaf14cf2455
3
+ size 1112197096
optimizer.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:2c68f3ed8eaeb602a94a167d2f1cef89b9665019c918e50882f194581c2c1c48
3
- size 940212619
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0e2051e1304f9cdfe43309a1548502446d35d8bb0867f9968f7c9b2b18658dc1
3
+ size 2219789707
rng_state.pth CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:f3ca820c499203a7884efe42f8add1720c864beea265fa9eea32b7023634d03e
3
  size 14645
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:50e4a16b58969404282a68dcf95227f3a6dbfa441535e172c6a1f5d3ab248361
3
  size 14645
scheduler.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:d682d4736c51a8bddec1240460a5cab5736031f55035e379f81f659ab4d9c0bb
3
  size 1465
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8c2cc6b245aaf5dedc3da5a396ac5028b3d52e53a67e380ea18bc62dc40b3e4e
3
  size 1465
special_tokens_map.json CHANGED
@@ -22,7 +22,7 @@
22
  },
23
  "mask_token": {
24
  "content": "<mask>",
25
- "lstrip": false,
26
  "normalized": false,
27
  "rstrip": false,
28
  "single_word": false
 
22
  },
23
  "mask_token": {
24
  "content": "<mask>",
25
+ "lstrip": true,
26
  "normalized": false,
27
  "rstrip": false,
28
  "single_word": false
tokenizer.json CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:ef04f2b385d1514f500e779207ace0f53e30895ce37563179e29f4022d28ca38
3
- size 17083053
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:883b037111086fd4dfebbbc9b7cee11e1517b5e0c0514879478661440f137085
3
+ size 17082987
tokenizer_config.json CHANGED
@@ -34,7 +34,7 @@
34
  },
35
  "250001": {
36
  "content": "<mask>",
37
- "lstrip": false,
38
  "normalized": false,
39
  "rstrip": false,
40
  "single_word": false,
@@ -50,7 +50,6 @@
50
  "model_max_length": 512,
51
  "pad_token": "<pad>",
52
  "sep_token": "</s>",
53
- "sp_model_kwargs": {},
54
  "tokenizer_class": "XLMRobertaTokenizer",
55
  "unk_token": "<unk>"
56
  }
 
34
  },
35
  "250001": {
36
  "content": "<mask>",
37
+ "lstrip": true,
38
  "normalized": false,
39
  "rstrip": false,
40
  "single_word": false,
 
50
  "model_max_length": 512,
51
  "pad_token": "<pad>",
52
  "sep_token": "</s>",
 
53
  "tokenizer_class": "XLMRobertaTokenizer",
54
  "unk_token": "<unk>"
55
  }
trainer_state.json CHANGED
@@ -1,492 +1,968 @@
1
  {
2
- "best_global_step": 2035,
3
- "best_metric": 8.790307998657227,
4
- "best_model_checkpoint": "printing_press/author-paraphrase/models/intfloat/multilingual-e5-small/checkpoint-2035",
5
  "epoch": 3.0,
6
  "eval_steps": 500,
7
- "global_step": 6105,
8
  "is_hyper_param_search": false,
9
  "is_local_process_zero": true,
10
  "is_world_process_zero": true,
11
  "log_history": [
12
  {
13
- "epoch": 0.04914004914004914,
14
- "grad_norm": 6.021224498748779,
15
- "learning_rate": 4.9513513513513516e-05,
16
- "loss": 3.1392,
17
  "step": 100
18
  },
19
  {
20
- "epoch": 0.09828009828009827,
21
- "grad_norm": 5.465278625488281,
22
- "learning_rate": 4.902211302211302e-05,
23
- "loss": 2.8632,
24
  "step": 200
25
  },
26
  {
27
- "epoch": 0.14742014742014742,
28
- "grad_norm": 6.197441577911377,
29
- "learning_rate": 4.853071253071254e-05,
30
- "loss": 2.6963,
31
  "step": 300
32
  },
33
  {
34
- "epoch": 0.19656019656019655,
35
- "grad_norm": 5.761979579925537,
36
- "learning_rate": 4.803931203931204e-05,
37
- "loss": 2.6594,
38
  "step": 400
39
  },
40
  {
41
- "epoch": 0.2457002457002457,
42
- "grad_norm": 6.232619285583496,
43
- "learning_rate": 4.754791154791155e-05,
44
- "loss": 2.584,
45
  "step": 500
46
  },
47
  {
48
- "epoch": 0.29484029484029484,
49
- "grad_norm": 5.387997627258301,
50
- "learning_rate": 4.705651105651106e-05,
51
- "loss": 2.5429,
52
  "step": 600
53
  },
54
  {
55
- "epoch": 0.343980343980344,
56
- "grad_norm": 5.559869766235352,
57
- "learning_rate": 4.656511056511057e-05,
58
- "loss": 2.5105,
59
  "step": 700
60
  },
61
  {
62
- "epoch": 0.3931203931203931,
63
- "grad_norm": 6.779473781585693,
64
- "learning_rate": 4.6073710073710074e-05,
65
- "loss": 2.4482,
66
  "step": 800
67
  },
68
  {
69
- "epoch": 0.44226044226044225,
70
- "grad_norm": 5.715710639953613,
71
- "learning_rate": 4.558230958230959e-05,
72
- "loss": 2.442,
73
  "step": 900
74
  },
75
  {
76
- "epoch": 0.4914004914004914,
77
- "grad_norm": 5.751602649688721,
78
- "learning_rate": 4.5090909090909095e-05,
79
- "loss": 2.4386,
80
  "step": 1000
81
  },
82
  {
83
- "epoch": 0.5405405405405406,
84
- "grad_norm": 6.389184951782227,
85
- "learning_rate": 4.45995085995086e-05,
86
- "loss": 2.3912,
87
  "step": 1100
88
  },
89
  {
90
- "epoch": 0.5896805896805897,
91
- "grad_norm": 5.976380825042725,
92
- "learning_rate": 4.410810810810811e-05,
93
- "loss": 2.3778,
94
  "step": 1200
95
  },
96
  {
97
- "epoch": 0.6388206388206388,
98
- "grad_norm": 5.9267120361328125,
99
- "learning_rate": 4.361670761670762e-05,
100
- "loss": 2.3481,
101
  "step": 1300
102
  },
103
  {
104
- "epoch": 0.687960687960688,
105
- "grad_norm": 5.765600681304932,
106
- "learning_rate": 4.312530712530713e-05,
107
- "loss": 2.3325,
108
  "step": 1400
109
  },
110
  {
111
- "epoch": 0.7371007371007371,
112
- "grad_norm": 5.1934285163879395,
113
- "learning_rate": 4.263390663390663e-05,
114
- "loss": 2.3082,
115
  "step": 1500
116
  },
117
  {
118
- "epoch": 0.7862407862407862,
119
- "grad_norm": 5.619853973388672,
120
- "learning_rate": 4.2142506142506146e-05,
121
- "loss": 2.2942,
122
  "step": 1600
123
  },
124
  {
125
- "epoch": 0.8353808353808354,
126
- "grad_norm": 5.406186103820801,
127
- "learning_rate": 4.165110565110565e-05,
128
- "loss": 2.3115,
129
  "step": 1700
130
  },
131
  {
132
- "epoch": 0.8845208845208845,
133
- "grad_norm": 4.95747184753418,
134
- "learning_rate": 4.115970515970517e-05,
135
- "loss": 2.2769,
136
  "step": 1800
137
  },
138
  {
139
- "epoch": 0.9336609336609336,
140
- "grad_norm": 6.417170524597168,
141
- "learning_rate": 4.066830466830467e-05,
142
- "loss": 2.2402,
143
  "step": 1900
144
  },
145
  {
146
- "epoch": 0.9828009828009828,
147
- "grad_norm": 6.344349384307861,
148
- "learning_rate": 4.017690417690418e-05,
149
- "loss": 2.2372,
150
  "step": 2000
151
  },
152
  {
153
- "epoch": 1.0,
154
- "eval_cosine_accuracy": 0.9193994404320385,
155
- "eval_cosine_accuracy_threshold": 0.8741108775138855,
156
- "eval_cosine_ap": 0.8367251348799856,
157
- "eval_cosine_f1": 0.7428981208338088,
158
- "eval_cosine_f1_threshold": 0.8662768006324768,
159
- "eval_cosine_mcc": 0.6915836823172079,
160
- "eval_cosine_precision": 0.7443417485874784,
161
- "eval_cosine_recall": 0.741460081983213,
162
- "eval_loss": 8.790307998657227,
163
- "eval_runtime": 1127.8164,
164
- "eval_samples_per_second": 163.527,
165
- "eval_steps_per_second": 1.203,
166
- "step": 2035
167
- },
168
- {
169
- "epoch": 1.031941031941032,
170
- "grad_norm": 5.7056660652160645,
171
- "learning_rate": 3.968550368550369e-05,
172
- "loss": 2.1558,
173
  "step": 2100
174
  },
175
  {
176
- "epoch": 1.0810810810810811,
177
- "grad_norm": 6.209120273590088,
178
- "learning_rate": 3.91941031941032e-05,
179
- "loss": 2.0741,
180
  "step": 2200
181
  },
182
  {
183
- "epoch": 1.1302211302211302,
184
- "grad_norm": 6.506903648376465,
185
- "learning_rate": 3.8702702702702704e-05,
186
- "loss": 2.0783,
187
  "step": 2300
188
  },
189
  {
190
- "epoch": 1.1793611793611793,
191
- "grad_norm": 5.882590293884277,
192
- "learning_rate": 3.821130221130221e-05,
193
- "loss": 2.0596,
194
  "step": 2400
195
  },
196
  {
197
- "epoch": 1.2285012285012284,
198
- "grad_norm": 5.904178619384766,
199
- "learning_rate": 3.7719901719901725e-05,
200
- "loss": 2.0398,
201
  "step": 2500
202
  },
203
  {
204
- "epoch": 1.2776412776412776,
205
- "grad_norm": 6.1165924072265625,
206
- "learning_rate": 3.7228501228501226e-05,
207
- "loss": 2.0504,
208
  "step": 2600
209
  },
210
  {
211
- "epoch": 1.3267813267813269,
212
- "grad_norm": 5.996497631072998,
213
- "learning_rate": 3.673710073710074e-05,
214
- "loss": 2.0674,
215
  "step": 2700
216
  },
217
  {
218
- "epoch": 1.375921375921376,
219
- "grad_norm": 7.482200622558594,
220
- "learning_rate": 3.624570024570025e-05,
221
- "loss": 2.0305,
222
  "step": 2800
223
  },
224
  {
225
- "epoch": 1.425061425061425,
226
- "grad_norm": 5.874721527099609,
227
- "learning_rate": 3.575429975429976e-05,
228
- "loss": 2.0372,
229
  "step": 2900
230
  },
231
  {
232
- "epoch": 1.4742014742014742,
233
- "grad_norm": 6.226138114929199,
234
- "learning_rate": 3.526289926289926e-05,
235
- "loss": 2.0271,
236
  "step": 3000
237
  },
238
  {
239
- "epoch": 1.5233415233415233,
240
- "grad_norm": 6.130570411682129,
241
- "learning_rate": 3.4771498771498776e-05,
242
- "loss": 2.007,
243
  "step": 3100
244
  },
245
  {
246
- "epoch": 1.5724815724815726,
247
- "grad_norm": 6.400360584259033,
248
- "learning_rate": 3.428009828009828e-05,
249
- "loss": 2.0074,
250
  "step": 3200
251
  },
252
  {
253
- "epoch": 1.6216216216216215,
254
- "grad_norm": 6.5269341468811035,
255
- "learning_rate": 3.378869778869779e-05,
256
- "loss": 1.9978,
257
  "step": 3300
258
  },
259
  {
260
- "epoch": 1.6707616707616708,
261
- "grad_norm": 5.755340576171875,
262
- "learning_rate": 3.32972972972973e-05,
263
- "loss": 1.974,
264
  "step": 3400
265
  },
266
  {
267
- "epoch": 1.71990171990172,
268
- "grad_norm": 6.60841178894043,
269
- "learning_rate": 3.2805896805896805e-05,
270
- "loss": 1.9922,
271
  "step": 3500
272
  },
273
  {
274
- "epoch": 1.769041769041769,
275
- "grad_norm": 7.436590194702148,
276
- "learning_rate": 3.231449631449632e-05,
277
- "loss": 1.9743,
278
  "step": 3600
279
  },
280
  {
281
- "epoch": 1.8181818181818183,
282
- "grad_norm": 6.426479339599609,
283
- "learning_rate": 3.182309582309582e-05,
284
- "loss": 1.9536,
285
  "step": 3700
286
  },
287
  {
288
- "epoch": 1.8673218673218672,
289
- "grad_norm": 6.678687572479248,
290
- "learning_rate": 3.1331695331695334e-05,
291
- "loss": 1.9717,
292
  "step": 3800
293
  },
294
  {
295
- "epoch": 1.9164619164619165,
296
- "grad_norm": 6.789039134979248,
297
- "learning_rate": 3.084029484029484e-05,
298
- "loss": 1.9324,
299
  "step": 3900
300
  },
301
  {
302
- "epoch": 1.9656019656019657,
303
- "grad_norm": 6.285867691040039,
304
- "learning_rate": 3.0348894348894352e-05,
305
- "loss": 1.9275,
306
  "step": 4000
307
  },
308
  {
309
- "epoch": 2.0,
310
- "eval_cosine_accuracy": 0.9200772117032121,
311
- "eval_cosine_accuracy_threshold": 0.8600538969039917,
312
- "eval_cosine_ap": 0.8386338276944761,
313
- "eval_cosine_f1": 0.7448358673957842,
314
- "eval_cosine_f1_threshold": 0.8511393070220947,
315
- "eval_cosine_mcc": 0.6936559183741586,
316
- "eval_cosine_precision": 0.7430153128945579,
317
- "eval_cosine_recall": 0.7466653653458261,
318
- "eval_loss": 9.22366714477539,
319
- "eval_runtime": 1127.8509,
320
- "eval_samples_per_second": 163.522,
321
- "eval_steps_per_second": 1.203,
322
- "step": 4070
323
- },
324
- {
325
- "epoch": 2.0147420147420148,
326
- "grad_norm": 7.566407680511475,
327
- "learning_rate": 2.9857493857493856e-05,
328
- "loss": 1.9059,
329
  "step": 4100
330
  },
331
  {
332
- "epoch": 2.063882063882064,
333
- "grad_norm": 6.225943565368652,
334
- "learning_rate": 2.9366093366093367e-05,
335
- "loss": 1.7814,
336
  "step": 4200
337
  },
338
  {
339
- "epoch": 2.113022113022113,
340
- "grad_norm": 6.670715808868408,
341
- "learning_rate": 2.8874692874692877e-05,
342
- "loss": 1.7528,
343
  "step": 4300
344
  },
345
  {
346
- "epoch": 2.1621621621621623,
347
- "grad_norm": 8.022383689880371,
348
- "learning_rate": 2.8383292383292388e-05,
349
- "loss": 1.786,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
350
  "step": 4400
351
  },
352
  {
353
- "epoch": 2.211302211302211,
354
- "grad_norm": 7.271859169006348,
355
- "learning_rate": 2.7891891891891892e-05,
356
- "loss": 1.7963,
357
  "step": 4500
358
  },
359
  {
360
- "epoch": 2.2604422604422605,
361
- "grad_norm": 6.646120071411133,
362
- "learning_rate": 2.7400491400491403e-05,
363
- "loss": 1.7744,
364
  "step": 4600
365
  },
366
  {
367
- "epoch": 2.30958230958231,
368
- "grad_norm": 6.8744330406188965,
369
- "learning_rate": 2.6909090909090913e-05,
370
- "loss": 1.7753,
371
  "step": 4700
372
  },
373
  {
374
- "epoch": 2.3587223587223587,
375
- "grad_norm": 8.4681396484375,
376
- "learning_rate": 2.641769041769042e-05,
377
- "loss": 1.7671,
378
  "step": 4800
379
  },
380
  {
381
- "epoch": 2.407862407862408,
382
- "grad_norm": 6.876530170440674,
383
- "learning_rate": 2.5926289926289924e-05,
384
- "loss": 1.7832,
385
  "step": 4900
386
  },
387
  {
388
- "epoch": 2.457002457002457,
389
- "grad_norm": 6.952081680297852,
390
- "learning_rate": 2.5434889434889435e-05,
391
- "loss": 1.7715,
392
  "step": 5000
393
  },
394
  {
395
- "epoch": 2.506142506142506,
396
- "grad_norm": 7.127411365509033,
397
- "learning_rate": 2.4943488943488943e-05,
398
- "loss": 1.721,
399
  "step": 5100
400
  },
401
  {
402
- "epoch": 2.555282555282555,
403
- "grad_norm": 7.662739276885986,
404
- "learning_rate": 2.4452088452088453e-05,
405
- "loss": 1.7584,
406
  "step": 5200
407
  },
408
  {
409
- "epoch": 2.6044226044226044,
410
- "grad_norm": 7.068917751312256,
411
- "learning_rate": 2.396068796068796e-05,
412
- "loss": 1.7348,
413
  "step": 5300
414
  },
415
  {
416
- "epoch": 2.6535626535626538,
417
- "grad_norm": 6.970505714416504,
418
- "learning_rate": 2.346928746928747e-05,
419
- "loss": 1.7331,
420
  "step": 5400
421
  },
422
  {
423
- "epoch": 2.7027027027027026,
424
- "grad_norm": 7.2441182136535645,
425
- "learning_rate": 2.297788697788698e-05,
426
- "loss": 1.7274,
427
  "step": 5500
428
  },
429
  {
430
- "epoch": 2.751842751842752,
431
- "grad_norm": 7.026419162750244,
432
- "learning_rate": 2.248648648648649e-05,
433
- "loss": 1.7587,
434
  "step": 5600
435
  },
436
  {
437
- "epoch": 2.800982800982801,
438
- "grad_norm": 5.773149490356445,
439
- "learning_rate": 2.1995085995085997e-05,
440
- "loss": 1.7379,
441
  "step": 5700
442
  },
443
  {
444
- "epoch": 2.85012285012285,
445
- "grad_norm": 6.8897809982299805,
446
- "learning_rate": 2.1503685503685507e-05,
447
- "loss": 1.7579,
448
  "step": 5800
449
  },
450
  {
451
- "epoch": 2.899262899262899,
452
- "grad_norm": 7.6783766746521,
453
- "learning_rate": 2.1012285012285015e-05,
454
- "loss": 1.7573,
455
  "step": 5900
456
  },
457
  {
458
- "epoch": 2.9484029484029484,
459
- "grad_norm": 6.818024158477783,
460
- "learning_rate": 2.0520884520884522e-05,
461
- "loss": 1.708,
462
  "step": 6000
463
  },
464
  {
465
- "epoch": 2.9975429975429977,
466
- "grad_norm": 7.6112542152404785,
467
- "learning_rate": 2.002948402948403e-05,
468
- "loss": 1.7159,
469
  "step": 6100
470
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
471
  {
472
  "epoch": 3.0,
473
- "eval_cosine_accuracy": 0.9208363155269265,
474
- "eval_cosine_accuracy_threshold": 0.8324185609817505,
475
- "eval_cosine_ap": 0.8415735890036254,
476
- "eval_cosine_f1": 0.7493691493691494,
477
- "eval_cosine_f1_threshold": 0.8252550363540649,
478
- "eval_cosine_mcc": 0.6992932604835177,
479
- "eval_cosine_precision": 0.7499918532277512,
480
- "eval_cosine_recall": 0.7487474786908712,
481
- "eval_loss": 9.839451789855957,
482
- "eval_runtime": 1129.5778,
483
- "eval_samples_per_second": 163.272,
484
- "eval_steps_per_second": 1.201,
485
- "step": 6105
486
  }
487
  ],
488
  "logging_steps": 100,
489
- "max_steps": 10175,
490
  "num_input_tokens_seen": 0,
491
  "num_train_epochs": 5,
492
  "save_steps": 500,
@@ -503,7 +979,7 @@
503
  }
504
  },
505
  "total_flos": 0.0,
506
- "train_batch_size": 136,
507
  "trial_name": null,
508
  "trial_params": null
509
  }
 
1
  {
2
+ "best_global_step": 4324,
3
+ "best_metric": 8.351966857910156,
4
+ "best_model_checkpoint": "printing_press/author-paraphrase/models/intfloat/multilingual-e5-base/checkpoint-4324",
5
  "epoch": 3.0,
6
  "eval_steps": 500,
7
+ "global_step": 12972,
8
  "is_hyper_param_search": false,
9
  "is_local_process_zero": true,
10
  "is_world_process_zero": true,
11
  "log_history": [
12
  {
13
+ "epoch": 0.02312673450508788,
14
+ "grad_norm": 10.87702751159668,
15
+ "learning_rate": 4.977104532839963e-05,
16
+ "loss": 2.4824,
17
  "step": 100
18
  },
19
  {
20
+ "epoch": 0.04625346901017576,
21
+ "grad_norm": 11.214370727539062,
22
+ "learning_rate": 4.953977798334876e-05,
23
+ "loss": 2.2607,
24
  "step": 200
25
  },
26
  {
27
+ "epoch": 0.06938020351526364,
28
+ "grad_norm": 9.89275074005127,
29
+ "learning_rate": 4.930851063829787e-05,
30
+ "loss": 2.1716,
31
  "step": 300
32
  },
33
  {
34
+ "epoch": 0.09250693802035152,
35
+ "grad_norm": 8.298206329345703,
36
+ "learning_rate": 4.9077243293247e-05,
37
+ "loss": 2.0981,
38
  "step": 400
39
  },
40
  {
41
+ "epoch": 0.11563367252543941,
42
+ "grad_norm": 9.424484252929688,
43
+ "learning_rate": 4.8845975948196116e-05,
44
+ "loss": 1.9617,
45
  "step": 500
46
  },
47
  {
48
+ "epoch": 0.13876040703052728,
49
+ "grad_norm": 8.752665519714355,
50
+ "learning_rate": 4.8614708603145235e-05,
51
+ "loss": 1.987,
52
  "step": 600
53
  },
54
  {
55
+ "epoch": 0.16188714153561518,
56
+ "grad_norm": 9.41073989868164,
57
+ "learning_rate": 4.838344125809436e-05,
58
+ "loss": 1.9429,
59
  "step": 700
60
  },
61
  {
62
+ "epoch": 0.18501387604070305,
63
+ "grad_norm": 11.025175094604492,
64
+ "learning_rate": 4.815217391304348e-05,
65
+ "loss": 1.9398,
66
  "step": 800
67
  },
68
  {
69
+ "epoch": 0.20814061054579094,
70
+ "grad_norm": 9.1474027633667,
71
+ "learning_rate": 4.79209065679926e-05,
72
+ "loss": 1.8745,
73
  "step": 900
74
  },
75
  {
76
+ "epoch": 0.23126734505087881,
77
+ "grad_norm": 8.528326034545898,
78
+ "learning_rate": 4.7689639222941726e-05,
79
+ "loss": 1.8484,
80
  "step": 1000
81
  },
82
  {
83
+ "epoch": 0.2543940795559667,
84
+ "grad_norm": 9.510236740112305,
85
+ "learning_rate": 4.7458371877890846e-05,
86
+ "loss": 1.846,
87
  "step": 1100
88
  },
89
  {
90
+ "epoch": 0.27752081406105455,
91
+ "grad_norm": 10.6268892288208,
92
+ "learning_rate": 4.7227104532839965e-05,
93
+ "loss": 1.7953,
94
  "step": 1200
95
  },
96
  {
97
+ "epoch": 0.30064754856614245,
98
+ "grad_norm": 9.342752456665039,
99
+ "learning_rate": 4.699583718778909e-05,
100
+ "loss": 1.8189,
101
  "step": 1300
102
  },
103
  {
104
+ "epoch": 0.32377428307123035,
105
+ "grad_norm": 8.43892765045166,
106
+ "learning_rate": 4.676456984273821e-05,
107
+ "loss": 1.7775,
108
  "step": 1400
109
  },
110
  {
111
+ "epoch": 0.34690101757631825,
112
+ "grad_norm": 12.730334281921387,
113
+ "learning_rate": 4.653330249768733e-05,
114
+ "loss": 1.76,
115
  "step": 1500
116
  },
117
  {
118
+ "epoch": 0.3700277520814061,
119
+ "grad_norm": 27.328882217407227,
120
+ "learning_rate": 4.6302035152636456e-05,
121
+ "loss": 1.7638,
122
  "step": 1600
123
  },
124
  {
125
+ "epoch": 0.393154486586494,
126
+ "grad_norm": 9.745387077331543,
127
+ "learning_rate": 4.607076780758557e-05,
128
+ "loss": 1.7039,
129
  "step": 1700
130
  },
131
  {
132
+ "epoch": 0.4162812210915819,
133
+ "grad_norm": 9.242189407348633,
134
+ "learning_rate": 4.583950046253469e-05,
135
+ "loss": 1.706,
136
  "step": 1800
137
  },
138
  {
139
+ "epoch": 0.43940795559666973,
140
+ "grad_norm": 9.079887390136719,
141
+ "learning_rate": 4.5608233117483814e-05,
142
+ "loss": 1.7255,
143
  "step": 1900
144
  },
145
  {
146
+ "epoch": 0.46253469010175763,
147
+ "grad_norm": 8.81753158569336,
148
+ "learning_rate": 4.537696577243293e-05,
149
+ "loss": 1.705,
150
  "step": 2000
151
  },
152
  {
153
+ "epoch": 0.4856614246068455,
154
+ "grad_norm": 10.990808486938477,
155
+ "learning_rate": 4.514569842738205e-05,
156
+ "loss": 1.6823,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
157
  "step": 2100
158
  },
159
  {
160
+ "epoch": 0.5087881591119334,
161
+ "grad_norm": 9.342223167419434,
162
+ "learning_rate": 4.491443108233118e-05,
163
+ "loss": 1.6921,
164
  "step": 2200
165
  },
166
  {
167
+ "epoch": 0.5319148936170213,
168
+ "grad_norm": 9.898025512695312,
169
+ "learning_rate": 4.46831637372803e-05,
170
+ "loss": 1.6801,
171
  "step": 2300
172
  },
173
  {
174
+ "epoch": 0.5550416281221091,
175
+ "grad_norm": 8.3154296875,
176
+ "learning_rate": 4.445189639222942e-05,
177
+ "loss": 1.6547,
178
  "step": 2400
179
  },
180
  {
181
+ "epoch": 0.5781683626271971,
182
+ "grad_norm": 8.5817232131958,
183
+ "learning_rate": 4.422062904717854e-05,
184
+ "loss": 1.6512,
185
  "step": 2500
186
  },
187
  {
188
+ "epoch": 0.6012950971322849,
189
+ "grad_norm": 8.931965827941895,
190
+ "learning_rate": 4.398936170212766e-05,
191
+ "loss": 1.6466,
192
  "step": 2600
193
  },
194
  {
195
+ "epoch": 0.6244218316373727,
196
+ "grad_norm": 10.271759033203125,
197
+ "learning_rate": 4.375809435707678e-05,
198
+ "loss": 1.6545,
199
  "step": 2700
200
  },
201
  {
202
+ "epoch": 0.6475485661424607,
203
+ "grad_norm": 8.621456146240234,
204
+ "learning_rate": 4.352682701202591e-05,
205
+ "loss": 1.5985,
206
  "step": 2800
207
  },
208
  {
209
+ "epoch": 0.6706753006475485,
210
+ "grad_norm": 9.487942695617676,
211
+ "learning_rate": 4.329555966697503e-05,
212
+ "loss": 1.5941,
213
  "step": 2900
214
  },
215
  {
216
+ "epoch": 0.6938020351526365,
217
+ "grad_norm": 9.156566619873047,
218
+ "learning_rate": 4.306429232192415e-05,
219
+ "loss": 1.6178,
220
  "step": 3000
221
  },
222
  {
223
+ "epoch": 0.7169287696577243,
224
+ "grad_norm": 11.126986503601074,
225
+ "learning_rate": 4.2833024976873266e-05,
226
+ "loss": 1.6035,
227
  "step": 3100
228
  },
229
  {
230
+ "epoch": 0.7400555041628122,
231
+ "grad_norm": 10.304364204406738,
232
+ "learning_rate": 4.2601757631822385e-05,
233
+ "loss": 1.568,
234
  "step": 3200
235
  },
236
  {
237
+ "epoch": 0.7631822386679001,
238
+ "grad_norm": 9.863882064819336,
239
+ "learning_rate": 4.2370490286771505e-05,
240
+ "loss": 1.5733,
241
  "step": 3300
242
  },
243
  {
244
+ "epoch": 0.786308973172988,
245
+ "grad_norm": 10.20065689086914,
246
+ "learning_rate": 4.213922294172063e-05,
247
+ "loss": 1.5841,
248
  "step": 3400
249
  },
250
  {
251
+ "epoch": 0.8094357076780758,
252
+ "grad_norm": 9.240017890930176,
253
+ "learning_rate": 4.190795559666975e-05,
254
+ "loss": 1.5869,
255
  "step": 3500
256
  },
257
  {
258
+ "epoch": 0.8325624421831638,
259
+ "grad_norm": 9.452214241027832,
260
+ "learning_rate": 4.167668825161887e-05,
261
+ "loss": 1.5806,
262
  "step": 3600
263
  },
264
  {
265
+ "epoch": 0.8556891766882516,
266
+ "grad_norm": 10.292157173156738,
267
+ "learning_rate": 4.1445420906567996e-05,
268
+ "loss": 1.5682,
269
  "step": 3700
270
  },
271
  {
272
+ "epoch": 0.8788159111933395,
273
+ "grad_norm": 12.675426483154297,
274
+ "learning_rate": 4.1214153561517115e-05,
275
+ "loss": 1.553,
276
  "step": 3800
277
  },
278
  {
279
+ "epoch": 0.9019426456984274,
280
+ "grad_norm": 11.213452339172363,
281
+ "learning_rate": 4.0982886216466234e-05,
282
+ "loss": 1.5564,
283
  "step": 3900
284
  },
285
  {
286
+ "epoch": 0.9250693802035153,
287
+ "grad_norm": 11.688211441040039,
288
+ "learning_rate": 4.075161887141536e-05,
289
+ "loss": 1.5389,
290
  "step": 4000
291
  },
292
  {
293
+ "epoch": 0.9481961147086031,
294
+ "grad_norm": 19.469762802124023,
295
+ "learning_rate": 4.052035152636448e-05,
296
+ "loss": 1.5091,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
297
  "step": 4100
298
  },
299
  {
300
+ "epoch": 0.971322849213691,
301
+ "grad_norm": 11.896566390991211,
302
+ "learning_rate": 4.02890841813136e-05,
303
+ "loss": 1.5358,
304
  "step": 4200
305
  },
306
  {
307
+ "epoch": 0.9944495837187789,
308
+ "grad_norm": 26.442358016967773,
309
+ "learning_rate": 4.0057816836262725e-05,
310
+ "loss": 1.512,
311
  "step": 4300
312
  },
313
  {
314
+ "epoch": 1.0,
315
+ "eval_cosine_accuracy": 0.9215032424577613,
316
+ "eval_cosine_accuracy_threshold": 0.8687731027603149,
317
+ "eval_cosine_ap": 0.844612185583288,
318
+ "eval_cosine_f1": 0.7527066450567261,
319
+ "eval_cosine_f1_threshold": 0.8616929054260254,
320
+ "eval_cosine_mcc": 0.7030390290848069,
321
+ "eval_cosine_precision": 0.7499838511724048,
322
+ "eval_cosine_recall": 0.7554492810202356,
323
+ "eval_loss": 8.351966857910156,
324
+ "eval_runtime": 1467.219,
325
+ "eval_samples_per_second": 125.699,
326
+ "eval_steps_per_second": 1.964,
327
+ "step": 4324
328
+ },
329
+ {
330
+ "epoch": 1.0175763182238668,
331
+ "grad_norm": 10.624114990234375,
332
+ "learning_rate": 3.9826549491211844e-05,
333
+ "loss": 1.4186,
334
  "step": 4400
335
  },
336
  {
337
+ "epoch": 1.0407030527289547,
338
+ "grad_norm": 10.404448509216309,
339
+ "learning_rate": 3.9595282146160964e-05,
340
+ "loss": 1.4075,
341
  "step": 4500
342
  },
343
  {
344
+ "epoch": 1.0638297872340425,
345
+ "grad_norm": 9.519079208374023,
346
+ "learning_rate": 3.936401480111008e-05,
347
+ "loss": 1.3934,
348
  "step": 4600
349
  },
350
  {
351
+ "epoch": 1.0869565217391304,
352
+ "grad_norm": 9.599648475646973,
353
+ "learning_rate": 3.91327474560592e-05,
354
+ "loss": 1.3799,
355
  "step": 4700
356
  },
357
  {
358
+ "epoch": 1.1100832562442182,
359
+ "grad_norm": 14.585043907165527,
360
+ "learning_rate": 3.890148011100833e-05,
361
+ "loss": 1.3597,
362
  "step": 4800
363
  },
364
  {
365
+ "epoch": 1.1332099907493063,
366
+ "grad_norm": 11.30241584777832,
367
+ "learning_rate": 3.867021276595745e-05,
368
+ "loss": 1.3351,
369
  "step": 4900
370
  },
371
  {
372
+ "epoch": 1.1563367252543941,
373
+ "grad_norm": 10.888731002807617,
374
+ "learning_rate": 3.843894542090657e-05,
375
+ "loss": 1.3082,
376
  "step": 5000
377
  },
378
  {
379
+ "epoch": 1.179463459759482,
380
+ "grad_norm": 9.09694766998291,
381
+ "learning_rate": 3.820767807585569e-05,
382
+ "loss": 1.3105,
383
  "step": 5100
384
  },
385
  {
386
+ "epoch": 1.2025901942645698,
387
+ "grad_norm": 13.162363052368164,
388
+ "learning_rate": 3.797641073080481e-05,
389
+ "loss": 1.2948,
390
  "step": 5200
391
  },
392
  {
393
+ "epoch": 1.2257169287696577,
394
+ "grad_norm": 26.262020111083984,
395
+ "learning_rate": 3.774514338575393e-05,
396
+ "loss": 1.3486,
397
  "step": 5300
398
  },
399
  {
400
+ "epoch": 1.2488436632747457,
401
+ "grad_norm": 21.114444732666016,
402
+ "learning_rate": 3.751387604070306e-05,
403
+ "loss": 1.3155,
404
  "step": 5400
405
  },
406
  {
407
+ "epoch": 1.2719703977798336,
408
+ "grad_norm": 10.668533325195312,
409
+ "learning_rate": 3.728260869565218e-05,
410
+ "loss": 1.2761,
411
  "step": 5500
412
  },
413
  {
414
+ "epoch": 1.2950971322849214,
415
+ "grad_norm": 12.978555679321289,
416
+ "learning_rate": 3.70513413506013e-05,
417
+ "loss": 1.2541,
418
  "step": 5600
419
  },
420
  {
421
+ "epoch": 1.3182238667900092,
422
+ "grad_norm": 13.066253662109375,
423
+ "learning_rate": 3.682007400555042e-05,
424
+ "loss": 1.2346,
425
  "step": 5700
426
  },
427
  {
428
+ "epoch": 1.341350601295097,
429
+ "grad_norm": 11.974534034729004,
430
+ "learning_rate": 3.658880666049954e-05,
431
+ "loss": 1.2285,
432
  "step": 5800
433
  },
434
  {
435
+ "epoch": 1.364477335800185,
436
+ "grad_norm": 12.733871459960938,
437
+ "learning_rate": 3.6357539315448655e-05,
438
+ "loss": 1.2013,
439
  "step": 5900
440
  },
441
  {
442
+ "epoch": 1.3876040703052728,
443
+ "grad_norm": 13.184294700622559,
444
+ "learning_rate": 3.612627197039778e-05,
445
+ "loss": 1.1986,
446
  "step": 6000
447
  },
448
  {
449
+ "epoch": 1.4107308048103608,
450
+ "grad_norm": 10.159133911132812,
451
+ "learning_rate": 3.58950046253469e-05,
452
+ "loss": 1.1755,
453
  "step": 6100
454
  },
455
+ {
456
+ "epoch": 1.4338575393154487,
457
+ "grad_norm": 12.212769508361816,
458
+ "learning_rate": 3.566373728029602e-05,
459
+ "loss": 1.1937,
460
+ "step": 6200
461
+ },
462
+ {
463
+ "epoch": 1.4569842738205365,
464
+ "grad_norm": 15.231562614440918,
465
+ "learning_rate": 3.5432469935245146e-05,
466
+ "loss": 1.202,
467
+ "step": 6300
468
+ },
469
+ {
470
+ "epoch": 1.4801110083256244,
471
+ "grad_norm": 14.845349311828613,
472
+ "learning_rate": 3.5201202590194265e-05,
473
+ "loss": 1.1607,
474
+ "step": 6400
475
+ },
476
+ {
477
+ "epoch": 1.5032377428307124,
478
+ "grad_norm": 12.119277954101562,
479
+ "learning_rate": 3.4969935245143384e-05,
480
+ "loss": 1.2116,
481
+ "step": 6500
482
+ },
483
+ {
484
+ "epoch": 1.5263644773358003,
485
+ "grad_norm": 12.19071102142334,
486
+ "learning_rate": 3.473866790009251e-05,
487
+ "loss": 1.1797,
488
+ "step": 6600
489
+ },
490
+ {
491
+ "epoch": 1.5494912118408881,
492
+ "grad_norm": 12.373006820678711,
493
+ "learning_rate": 3.450740055504163e-05,
494
+ "loss": 1.1571,
495
+ "step": 6700
496
+ },
497
+ {
498
+ "epoch": 1.572617946345976,
499
+ "grad_norm": 14.413765907287598,
500
+ "learning_rate": 3.427613320999075e-05,
501
+ "loss": 1.1526,
502
+ "step": 6800
503
+ },
504
+ {
505
+ "epoch": 1.5957446808510638,
506
+ "grad_norm": 12.39973258972168,
507
+ "learning_rate": 3.4044865864939875e-05,
508
+ "loss": 1.1438,
509
+ "step": 6900
510
+ },
511
+ {
512
+ "epoch": 1.6188714153561516,
513
+ "grad_norm": 15.030281066894531,
514
+ "learning_rate": 3.3813598519888994e-05,
515
+ "loss": 1.1634,
516
+ "step": 7000
517
+ },
518
+ {
519
+ "epoch": 1.6419981498612395,
520
+ "grad_norm": 15.445487022399902,
521
+ "learning_rate": 3.3582331174838114e-05,
522
+ "loss": 1.1367,
523
+ "step": 7100
524
+ },
525
+ {
526
+ "epoch": 1.6651248843663273,
527
+ "grad_norm": 15.726140022277832,
528
+ "learning_rate": 3.335106382978724e-05,
529
+ "loss": 1.1133,
530
+ "step": 7200
531
+ },
532
+ {
533
+ "epoch": 1.6882516188714154,
534
+ "grad_norm": 12.940756797790527,
535
+ "learning_rate": 3.311979648473636e-05,
536
+ "loss": 1.1156,
537
+ "step": 7300
538
+ },
539
+ {
540
+ "epoch": 1.7113783533765032,
541
+ "grad_norm": 14.717055320739746,
542
+ "learning_rate": 3.288852913968548e-05,
543
+ "loss": 1.1102,
544
+ "step": 7400
545
+ },
546
+ {
547
+ "epoch": 1.734505087881591,
548
+ "grad_norm": 12.775771141052246,
549
+ "learning_rate": 3.26572617946346e-05,
550
+ "loss": 1.1123,
551
+ "step": 7500
552
+ },
553
+ {
554
+ "epoch": 1.7576318223866791,
555
+ "grad_norm": 12.529901504516602,
556
+ "learning_rate": 3.242599444958372e-05,
557
+ "loss": 1.1066,
558
+ "step": 7600
559
+ },
560
+ {
561
+ "epoch": 1.780758556891767,
562
+ "grad_norm": 12.506126403808594,
563
+ "learning_rate": 3.219472710453284e-05,
564
+ "loss": 1.1291,
565
+ "step": 7700
566
+ },
567
+ {
568
+ "epoch": 1.8038852913968548,
569
+ "grad_norm": 16.9326114654541,
570
+ "learning_rate": 3.196345975948196e-05,
571
+ "loss": 1.1094,
572
+ "step": 7800
573
+ },
574
+ {
575
+ "epoch": 1.8270120259019427,
576
+ "grad_norm": 12.257229804992676,
577
+ "learning_rate": 3.173219241443108e-05,
578
+ "loss": 1.094,
579
+ "step": 7900
580
+ },
581
+ {
582
+ "epoch": 1.8501387604070305,
583
+ "grad_norm": 15.346363067626953,
584
+ "learning_rate": 3.150092506938021e-05,
585
+ "loss": 1.1585,
586
+ "step": 8000
587
+ },
588
+ {
589
+ "epoch": 1.8732654949121184,
590
+ "grad_norm": 12.800529479980469,
591
+ "learning_rate": 3.126965772432933e-05,
592
+ "loss": 1.077,
593
+ "step": 8100
594
+ },
595
+ {
596
+ "epoch": 1.8963922294172062,
597
+ "grad_norm": 14.989642143249512,
598
+ "learning_rate": 3.103839037927845e-05,
599
+ "loss": 1.108,
600
+ "step": 8200
601
+ },
602
+ {
603
+ "epoch": 1.919518963922294,
604
+ "grad_norm": 14.930624008178711,
605
+ "learning_rate": 3.080712303422757e-05,
606
+ "loss": 1.1431,
607
+ "step": 8300
608
+ },
609
+ {
610
+ "epoch": 1.942645698427382,
611
+ "grad_norm": 13.488320350646973,
612
+ "learning_rate": 3.057585568917669e-05,
613
+ "loss": 1.0784,
614
+ "step": 8400
615
+ },
616
+ {
617
+ "epoch": 1.96577243293247,
618
+ "grad_norm": 12.30305004119873,
619
+ "learning_rate": 3.034458834412581e-05,
620
+ "loss": 1.0834,
621
+ "step": 8500
622
+ },
623
+ {
624
+ "epoch": 1.9888991674375578,
625
+ "grad_norm": 13.152817726135254,
626
+ "learning_rate": 3.0113320999074934e-05,
627
+ "loss": 1.1268,
628
+ "step": 8600
629
+ },
630
+ {
631
+ "epoch": 2.0,
632
+ "eval_cosine_accuracy": 0.9209068037391286,
633
+ "eval_cosine_accuracy_threshold": 0.8271753191947937,
634
+ "eval_cosine_ap": 0.8450278341834471,
635
+ "eval_cosine_f1": 0.7521474761123443,
636
+ "eval_cosine_f1_threshold": 0.8189181089401245,
637
+ "eval_cosine_mcc": 0.701977809831327,
638
+ "eval_cosine_precision": 0.7438907980145093,
639
+ "eval_cosine_recall": 0.7605894983408159,
640
+ "eval_loss": 9.69921588897705,
641
+ "eval_runtime": 1472.4233,
642
+ "eval_samples_per_second": 125.255,
643
+ "eval_steps_per_second": 1.957,
644
+ "step": 8648
645
+ },
646
+ {
647
+ "epoch": 2.012025901942646,
648
+ "grad_norm": 15.71314811706543,
649
+ "learning_rate": 2.9882053654024057e-05,
650
+ "loss": 1.0443,
651
+ "step": 8700
652
+ },
653
+ {
654
+ "epoch": 2.0351526364477337,
655
+ "grad_norm": 14.261523246765137,
656
+ "learning_rate": 2.9650786308973173e-05,
657
+ "loss": 0.9715,
658
+ "step": 8800
659
+ },
660
+ {
661
+ "epoch": 2.0582793709528215,
662
+ "grad_norm": 13.405384063720703,
663
+ "learning_rate": 2.9419518963922292e-05,
664
+ "loss": 0.957,
665
+ "step": 8900
666
+ },
667
+ {
668
+ "epoch": 2.0814061054579094,
669
+ "grad_norm": 13.5853853225708,
670
+ "learning_rate": 2.9188251618871415e-05,
671
+ "loss": 0.9784,
672
+ "step": 9000
673
+ },
674
+ {
675
+ "epoch": 2.1045328399629972,
676
+ "grad_norm": 14.918572425842285,
677
+ "learning_rate": 2.8956984273820538e-05,
678
+ "loss": 0.9581,
679
+ "step": 9100
680
+ },
681
+ {
682
+ "epoch": 2.127659574468085,
683
+ "grad_norm": 13.354079246520996,
684
+ "learning_rate": 2.8725716928769657e-05,
685
+ "loss": 0.9569,
686
+ "step": 9200
687
+ },
688
+ {
689
+ "epoch": 2.150786308973173,
690
+ "grad_norm": 15.024975776672363,
691
+ "learning_rate": 2.849444958371878e-05,
692
+ "loss": 0.9518,
693
+ "step": 9300
694
+ },
695
+ {
696
+ "epoch": 2.1739130434782608,
697
+ "grad_norm": 13.5723876953125,
698
+ "learning_rate": 2.8263182238667902e-05,
699
+ "loss": 0.9485,
700
+ "step": 9400
701
+ },
702
+ {
703
+ "epoch": 2.1970397779833486,
704
+ "grad_norm": 12.383338928222656,
705
+ "learning_rate": 2.8031914893617022e-05,
706
+ "loss": 0.9433,
707
+ "step": 9500
708
+ },
709
+ {
710
+ "epoch": 2.2201665124884364,
711
+ "grad_norm": 13.54079532623291,
712
+ "learning_rate": 2.7800647548566144e-05,
713
+ "loss": 0.9392,
714
+ "step": 9600
715
+ },
716
+ {
717
+ "epoch": 2.2432932469935247,
718
+ "grad_norm": 15.325400352478027,
719
+ "learning_rate": 2.7569380203515267e-05,
720
+ "loss": 0.9248,
721
+ "step": 9700
722
+ },
723
+ {
724
+ "epoch": 2.2664199814986126,
725
+ "grad_norm": 15.651702880859375,
726
+ "learning_rate": 2.7338112858464387e-05,
727
+ "loss": 0.9105,
728
+ "step": 9800
729
+ },
730
+ {
731
+ "epoch": 2.2895467160037004,
732
+ "grad_norm": 15.196932792663574,
733
+ "learning_rate": 2.710684551341351e-05,
734
+ "loss": 0.9769,
735
+ "step": 9900
736
+ },
737
+ {
738
+ "epoch": 2.3126734505087883,
739
+ "grad_norm": 13.820411682128906,
740
+ "learning_rate": 2.6875578168362632e-05,
741
+ "loss": 0.9502,
742
+ "step": 10000
743
+ },
744
+ {
745
+ "epoch": 2.335800185013876,
746
+ "grad_norm": 15.181200981140137,
747
+ "learning_rate": 2.664431082331175e-05,
748
+ "loss": 0.9604,
749
+ "step": 10100
750
+ },
751
+ {
752
+ "epoch": 2.358926919518964,
753
+ "grad_norm": 17.202287673950195,
754
+ "learning_rate": 2.6413043478260867e-05,
755
+ "loss": 0.9291,
756
+ "step": 10200
757
+ },
758
+ {
759
+ "epoch": 2.382053654024052,
760
+ "grad_norm": 16.69881248474121,
761
+ "learning_rate": 2.618177613320999e-05,
762
+ "loss": 0.9552,
763
+ "step": 10300
764
+ },
765
+ {
766
+ "epoch": 2.4051803885291396,
767
+ "grad_norm": 16.424978256225586,
768
+ "learning_rate": 2.5950508788159113e-05,
769
+ "loss": 0.9621,
770
+ "step": 10400
771
+ },
772
+ {
773
+ "epoch": 2.4283071230342275,
774
+ "grad_norm": 17.192264556884766,
775
+ "learning_rate": 2.5719241443108232e-05,
776
+ "loss": 0.9357,
777
+ "step": 10500
778
+ },
779
+ {
780
+ "epoch": 2.4514338575393153,
781
+ "grad_norm": 16.9178524017334,
782
+ "learning_rate": 2.5487974098057355e-05,
783
+ "loss": 0.9323,
784
+ "step": 10600
785
+ },
786
+ {
787
+ "epoch": 2.474560592044403,
788
+ "grad_norm": 11.07499885559082,
789
+ "learning_rate": 2.5256706753006477e-05,
790
+ "loss": 0.9327,
791
+ "step": 10700
792
+ },
793
+ {
794
+ "epoch": 2.4976873265494914,
795
+ "grad_norm": 28.121694564819336,
796
+ "learning_rate": 2.5025439407955597e-05,
797
+ "loss": 0.9067,
798
+ "step": 10800
799
+ },
800
+ {
801
+ "epoch": 2.520814061054579,
802
+ "grad_norm": 15.262548446655273,
803
+ "learning_rate": 2.479417206290472e-05,
804
+ "loss": 0.9411,
805
+ "step": 10900
806
+ },
807
+ {
808
+ "epoch": 2.543940795559667,
809
+ "grad_norm": 16.59368133544922,
810
+ "learning_rate": 2.4562904717853842e-05,
811
+ "loss": 0.9305,
812
+ "step": 11000
813
+ },
814
+ {
815
+ "epoch": 2.567067530064755,
816
+ "grad_norm": 12.39183235168457,
817
+ "learning_rate": 2.433163737280296e-05,
818
+ "loss": 0.9378,
819
+ "step": 11100
820
+ },
821
+ {
822
+ "epoch": 2.590194264569843,
823
+ "grad_norm": 19.034120559692383,
824
+ "learning_rate": 2.4100370027752084e-05,
825
+ "loss": 0.9171,
826
+ "step": 11200
827
+ },
828
+ {
829
+ "epoch": 2.6133209990749307,
830
+ "grad_norm": 16.727380752563477,
831
+ "learning_rate": 2.3869102682701204e-05,
832
+ "loss": 0.9074,
833
+ "step": 11300
834
+ },
835
+ {
836
+ "epoch": 2.6364477335800185,
837
+ "grad_norm": 17.257230758666992,
838
+ "learning_rate": 2.3637835337650323e-05,
839
+ "loss": 0.9262,
840
+ "step": 11400
841
+ },
842
+ {
843
+ "epoch": 2.6595744680851063,
844
+ "grad_norm": 18.956735610961914,
845
+ "learning_rate": 2.3406567992599446e-05,
846
+ "loss": 0.9063,
847
+ "step": 11500
848
+ },
849
+ {
850
+ "epoch": 2.682701202590194,
851
+ "grad_norm": 15.463052749633789,
852
+ "learning_rate": 2.317530064754857e-05,
853
+ "loss": 0.8814,
854
+ "step": 11600
855
+ },
856
+ {
857
+ "epoch": 2.705827937095282,
858
+ "grad_norm": 13.263724327087402,
859
+ "learning_rate": 2.2944033302497688e-05,
860
+ "loss": 0.9089,
861
+ "step": 11700
862
+ },
863
+ {
864
+ "epoch": 2.72895467160037,
865
+ "grad_norm": 15.596953392028809,
866
+ "learning_rate": 2.271276595744681e-05,
867
+ "loss": 0.9048,
868
+ "step": 11800
869
+ },
870
+ {
871
+ "epoch": 2.752081406105458,
872
+ "grad_norm": 16.61760711669922,
873
+ "learning_rate": 2.2481498612395933e-05,
874
+ "loss": 0.9268,
875
+ "step": 11900
876
+ },
877
+ {
878
+ "epoch": 2.7752081406105455,
879
+ "grad_norm": 16.409652709960938,
880
+ "learning_rate": 2.2250231267345052e-05,
881
+ "loss": 0.8913,
882
+ "step": 12000
883
+ },
884
+ {
885
+ "epoch": 2.798334875115634,
886
+ "grad_norm": 14.509052276611328,
887
+ "learning_rate": 2.2018963922294172e-05,
888
+ "loss": 0.9064,
889
+ "step": 12100
890
+ },
891
+ {
892
+ "epoch": 2.8214616096207217,
893
+ "grad_norm": 13.914097785949707,
894
+ "learning_rate": 2.1787696577243295e-05,
895
+ "loss": 0.8585,
896
+ "step": 12200
897
+ },
898
+ {
899
+ "epoch": 2.8445883441258095,
900
+ "grad_norm": 13.852704048156738,
901
+ "learning_rate": 2.1556429232192414e-05,
902
+ "loss": 0.878,
903
+ "step": 12300
904
+ },
905
+ {
906
+ "epoch": 2.8677150786308974,
907
+ "grad_norm": 15.329246520996094,
908
+ "learning_rate": 2.1325161887141537e-05,
909
+ "loss": 0.8612,
910
+ "step": 12400
911
+ },
912
+ {
913
+ "epoch": 2.890841813135985,
914
+ "grad_norm": 15.046784400939941,
915
+ "learning_rate": 2.109389454209066e-05,
916
+ "loss": 0.8799,
917
+ "step": 12500
918
+ },
919
+ {
920
+ "epoch": 2.913968547641073,
921
+ "grad_norm": 15.824640274047852,
922
+ "learning_rate": 2.086262719703978e-05,
923
+ "loss": 0.8541,
924
+ "step": 12600
925
+ },
926
+ {
927
+ "epoch": 2.937095282146161,
928
+ "grad_norm": 15.991775512695312,
929
+ "learning_rate": 2.0631359851988898e-05,
930
+ "loss": 0.8521,
931
+ "step": 12700
932
+ },
933
+ {
934
+ "epoch": 2.9602220166512487,
935
+ "grad_norm": 16.13051986694336,
936
+ "learning_rate": 2.040009250693802e-05,
937
+ "loss": 0.8582,
938
+ "step": 12800
939
+ },
940
+ {
941
+ "epoch": 2.9833487511563366,
942
+ "grad_norm": 14.818473815917969,
943
+ "learning_rate": 2.0168825161887143e-05,
944
+ "loss": 0.869,
945
+ "step": 12900
946
+ },
947
  {
948
  "epoch": 3.0,
949
+ "eval_cosine_accuracy": 0.9225280326197758,
950
+ "eval_cosine_accuracy_threshold": 0.7901061773300171,
951
+ "eval_cosine_ap": 0.8478615501518483,
952
+ "eval_cosine_f1": 0.7559554803436604,
953
+ "eval_cosine_f1_threshold": 0.7817596793174744,
954
+ "eval_cosine_mcc": 0.7071656901034916,
955
+ "eval_cosine_precision": 0.756201575623413,
956
+ "eval_cosine_recall": 0.7557095451883662,
957
+ "eval_loss": 10.411548614501953,
958
+ "eval_runtime": 1486.9605,
959
+ "eval_samples_per_second": 124.03,
960
+ "eval_steps_per_second": 1.938,
961
+ "step": 12972
962
  }
963
  ],
964
  "logging_steps": 100,
965
+ "max_steps": 21620,
966
  "num_input_tokens_seen": 0,
967
  "num_train_epochs": 5,
968
  "save_steps": 500,
 
979
  }
980
  },
981
  "total_flos": 0.0,
982
+ "train_batch_size": 64,
983
  "trial_name": null,
984
  "trial_params": null
985
  }
training_args.bin CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:5a9c5f531debf1f74c84f519c72b0b1c02f7f63a0b81d042e6d7aff896d6f1e0
3
  size 6097
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:05569020741fe60e0a967a0b69b87b45af12fa4c5d912666a335c2b0eea34235
3
  size 6097