--- language: - lv license: other license_name: research-use-mixed-upstream-data library_name: transformers pipeline_tag: summarization base_model: facebook/mbart-large-50-many-to-many-mmt base_model_relation: finetune tags: - summarization - mbart50 - low-resource - research - rahvusarhiiv datasets: - dennlinger/eur-lex-sum - FMISummarization/FMI_Summarization widget: - text: "Latvijas pētnieki paziņoja par jaunu projektu, kura mērķis ir uzlabot digitālo arhīvu pieejamību. Projekts paredz sadarboties ar bibliotēkām, universitātēm un valsts institūcijām." example_title: Latvian example model-index: - name: lv_summariser revision 2 results: - task: type: summarization name: Summarization dataset: name: Strict Wiki4 (wiki4_lv) type: wiki4_lv metrics: - name: ROUGE-L F1 type: rougeL value: 0.313275 - name: Reference content-token coverage type: reference_content_token_coverage value: 0.290934 - name: Prediction source-token precision type: prediction_source_content_token_precision value: 0.863554 - task: type: summarization name: Summarization dataset: name: Recent GDELT news (gdelt_lv) type: gdelt_lv metrics: - name: ROUGE-L F1 type: rougeL value: 0.201971 - name: Reference content-token coverage type: reference_content_token_coverage value: 0.217000 - name: Prediction source-token precision type: prediction_source_content_token_precision value: 0.813343 - task: type: summarization name: Summarization dataset: name: Held-out supported news (news_supported_lv) type: news_supported_lv metrics: - name: ROUGE-L F1 type: rougeL value: 0.284150 - name: Reference content-token coverage type: reference_content_token_coverage value: 0.283575 - name: Prediction source-token precision type: prediction_source_content_token_precision value: 0.816354 --- # Latvian summariser `Rahvusarhiiv/lv_summariser` is an experimental Latvian summarization model fine-tuned from `facebook/mbart-large-50-many-to-many-mmt`. The default branch contains revision 2, published in July 2026. The previously released checkpoint is preserved under the `v1` tag. Revision 2 improves supported-news and recent-news summarization, source grounding, and removal of learned EU legal boilerplate, while producing shorter encyclopedic summaries. This is a research release. It is not claimed to be state of the art; evaluate it on the target domain before production use. ## Intended Use - Summarization of Latvian news, legal/administrative, encyclopedic, or academic-style text. - Batch/offline summarization, evaluation, and retrieval-assisted workflows where outputs can be checked. - Not intended as a sole source of factual truth, legal advice, medical advice, or archival description without human review. ## How to Use ```python from transformers import AutoModelForSeq2SeqLM, AutoTokenizer repo_id = "Rahvusarhiiv/lv_summariser" language_code = "lv_LV" text = "Latvijas pētnieki paziņoja par jaunu projektu, kura mērķis ir uzlabot digitālo arhīvu pieejamību. Projekts paredz sadarboties ar bibliotēkām, universitātēm un valsts institūcijām." tokenizer = AutoTokenizer.from_pretrained(repo_id, src_lang=language_code) model = AutoModelForSeq2SeqLM.from_pretrained(repo_id) inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=768) language_token_id = tokenizer.convert_tokens_to_ids(language_code) summary_ids = model.generate( **inputs, forced_bos_token_id=language_token_id, max_new_tokens=512, num_beams=4, no_repeat_ngram_size=3, ) print(tokenizer.batch_decode(summary_ids, skip_special_tokens=True)[0]) ``` `convert_tokens_to_ids` is used deliberately. Current Transformers releases may load the tokenizer as `TokenizersBackend`, which does not expose the older tokenizer-specific `lang_code_to_id` attribute. ## Recommended Generation Settings - `max_input_tokens`: `768` - `max_new_tokens`: `512` - `num_beams`: `4` - `length_penalty`: `1.0` - `no_repeat_ngram_size`: `3` These are the settings used for the revision-2 comparison evaluation. ## Training Data Revision 2 was trained on 30,000 balanced examples after protected-benchmark exclusion: | family | training rows | share | |---|---:|---:| | Latvian news | 27,180 | 90.6% | | EU legal | 2,412 | 8.0% | | Academic/thesis | 408 | 1.4% | The sources include: - 27,143 accepted articles from the locally collected LSM archive and 37 additional LSM/Delfi RSS examples. - 2,412 cleaned `dennlinger/eur-lex-sum` examples. `FMISummarization/FMI_Summarization` was used to recover supporting source acts and select evidence; it was not added as duplicate sampling mass. - 408 Latvian academic OAI abstract/full-text pairs collected from Latvian university journal repositories. Legal targets underwent more aggressive summary-section extraction and boilerplate filtering. Long inputs use source-span evidence selection. The build excluded protected Wiki4, GDELT, and supported-news benchmark overlap using IDs, normalized URLs, exact source/summary/pair hashes, and shared token shingles. Training metadata: - Base model: `facebook/mbart-large-50-many-to-many-mmt` - Training rows: `30,000` - Validation rows: `708` - Max source tokens: `768` - Max target tokens: `384` - Seed: `20260713` ## Evaluation Automatic evaluation used 100 examples per benchmark with the same decoding settings for revision 1 and revision 2. | benchmark | ROUGE-L | ROUGE-1 | ROUGE-2 | ref coverage | src precision | words | repeat 4g | legalese | max hit | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:| | Strict Wiki4 | 0.313 | 0.347 | 0.251 | 0.291 | 0.864 | 28.7 | 0.000 | 0.000 | 0.000 | | Recent GDELT news | 0.202 | 0.240 | 0.104 | 0.217 | 0.813 | 21.0 | 0.000 | 0.000 | 0.000 | | Held-out supported news | 0.284 | 0.331 | 0.176 | 0.284 | 0.816 | 24.4 | 0.000 | 0.000 | 0.000 | ### Revision 1 Comparison | benchmark | v1 ROUGE-L | v2 ROUGE-L | v1 src precision | v2 src precision | |---|---:|---:|---:|---:| | Strict Wiki4 | 0.387 | 0.313 | 0.787 | 0.864 | | Recent GDELT news | 0.160 | 0.202 | 0.798 | 0.813 | | Held-out supported news | 0.154 | 0.284 | 0.779 | 0.816 | A local `gemma4:31b` Ollama judge evaluated five paired random examples per benchmark. This small comparative sample is not an absolute quality estimate. | benchmark | v1 factuality | v2 factuality | v1 coverage | v2 coverage | v1 overall | v2 overall | |---|---:|---:|---:|---:|---:|---:| | Strict Wiki4 | 3.20 | 3.80 | 2.20 | 2.20 | 2.60 | 2.60 | | Recent GDELT news | 3.00 | 2.80 | 1.80 | 2.80 | 2.00 | 2.40 | | Held-out supported news | 2.20 | 3.20 | 1.80 | 2.40 | 1.60 | 2.60 | Revision 2 was selected because it improved both automatic and judge results on the two news benchmarks, increased source precision in all three domains, and removed detected legalese artifacts. Wiki4 ROUGE and reference coverage declined because the new model produces shorter encyclopedic summaries; this is a known tradeoff rather than a universal improvement. ## Limitations and Failure Modes - The model can hallucinate or over-generalize details, especially when the input is long or under-specified. - Named entities, numbers, dates, and legal references should be verified against the source. - Revision 2 can under-summarize broad encyclopedic inputs and omit secondary points. - Summary length control was present in parts of training, but this release does not expose a guaranteed length-control interface. - Title-like opening phrases can still occur; the Wiki4 duplicate-title-prefix flag was `0.09` in this evaluation. - The training data mixes sources with different writing styles and upstream terms; users are responsible for checking suitability for their use case. ## Bias, Risks, and Ethical Considerations The model reflects the distributions and editorial choices of its source corpora. Its revision-2 mix is heavily weighted toward Latvian public-broadcaster news, so its style and topic distribution are not representative of all Latvian writing. It may reproduce biases present in news, legal, and academic sources. Human review is recommended for public-facing summaries and high-impact decisions. ## CO2 Emissions Training emissions were not instrumented with CodeCarbon or an equivalent tracker. No numeric CO2 estimate is reported. Training was local fine-tuning of an mBART-50 checkpoint rather than pretraining from scratch. ## Model Tree - Base model: `facebook/mbart-large-50-many-to-many-mmt` - Relationship: fine-tuned sequence-to-sequence summarization checkpoint. - Previous release: [`v1`](https://huggingface.co/Rahvusarhiiv/lv_summariser/tree/v1) ## License The model card uses `license: other` because the training mix contains multiple upstream data sources with different terms. Treat this as a research-use release and verify upstream data/license constraints before redistribution or commercial use.