# Pipeline de treino EDA no Hugging Face Documentacao canónica do fluxo. Repositórios e papéis: ver [`hf_registry.json`](../hf_registry.json). ## Visao geral ```mermaid flowchart LR subgraph dados DS[ml/datasets ou JSONL exportado] HFD[beAnalytic/eda-training-dataset] end subgraph treino SP[Space beAnalytic/Training] LORA[beAnalytic/eda-llm-qwen2.5-lora] end subgraph release MER[beAnalytic/eda-llm-qwen2.5-merged] CARD[beAnalytic/eda-llm-beAnalityc] end DS --> HFD HFD --> SP SP --> LORA LORA --> MER MER --> CARD ``` ## 1. Dados - **Fonte versionada no repo:** `ml/datasets/eda_training_dataset/` (Hugging Face `datasets` em disco). - **Não** versionar cache arrow dentro de `huggingface_training_config/` (está em `.gitignore`). - **Hub:** dataset de treino em `beAnalytic/eda-training-dataset` (JSONL por linha). - Upload (preferível: `DatasetDict` local em `ml/datasets/eda_training_dataset/`): ```bash cd ml/configs/huggingface_training_config export HF_TOKEN=... python scripts/push_dataset_to_hub.py ``` Alternativa: JSONL ou pasta com `python scripts/upload_dataset.py /caminho/real/para/arquivo.jsonl`. `push_dataset_to_hub.py` sem variável `DATASET_REPO` envia para `/eda-training-dataset`. Para `beAnalytic/eda-training-dataset` é preciso permissão na org e `export DATASET_REPO=beAnalytic/eda-training-dataset`. ## 2. Ficheiros do Space (Docker) O Space `beAnalytic/Training` precisa no mínimo: | Ficheiro | Função | |----------|--------| | `SPACE_README.md` | Copiar como `README.md` no clone do Space: frontmatter YAML (`sdk: docker`, `app_port`, etc.) | | `Dockerfile` | Imagem e `CMD` | | `app.py` | Entrada: invoca `train.py` | | `train.py` | LoRA + Trainer | | `requirements.txt` | Dependências | Copiar a partir desta pasta para o clone do Space: ```bash git clone https://huggingface.co/spaces/beAnalytic/Training cd Training cp /caminho/Be-Quick-Insights/ml/configs/huggingface_training_config/SPACE_README.md ./README.md cp /caminho/.../Dockerfile . cp /caminho/.../app.py . cp /caminho/.../train.py . cp /caminho/.../requirements.txt . git add README.md Dockerfile app.py train.py requirements.txt git commit -m "Atualizar treino EDA" git push ``` ## 3. Secrets do Space Settings: `https://huggingface.co/spaces/beAnalytic/Training/settings` → Repository secrets: | Variável | Valor típico | |----------|----------------| | `HF_TOKEN` | Token com permissão **write** (não commitar) | | `DATASET_REPO` | `beAnalytic/eda-training-dataset` | | `OUTPUT_REPO` | `beAnalytic/eda-llm-qwen2.5-lora` | | `MODEL_NAME` | `Qwen/Qwen2.5-1.5B-Instruct` | Sem `HF_TOKEN`, o treino corre mas o push de checkpoints para o Hub fica desligado; use checkpoints em `./results` localmente e upload manual se necessário. ## 4. Treino local ```bash cd ml/configs/huggingface_training_config export HF_TOKEN=... export DATASET_REPO=beAnalytic/eda-training-dataset export OUTPUT_REPO=beAnalytic/eda-llm-qwen2.5-lora python train.py ``` O [`Dockerfile`](../Dockerfile) usa `CMD ["python", "/app/app.py"]`, que apenas executa `train.py` (sem servidor TensorBoard no `app.py` atual; métricas em ficheiros de eventos, ver [METRICS.md](METRICS.md)). ### Upload manual de checkpoints Se `push_to_hub` estiver desligado ou falhar, envie a pasta de resultados (adapter, tokenizer, etc.) para o repo de modelo LoRA do registo: ```bash pip install "huggingface_hub>=1.0" hf auth login cd results bash ../scripts/clean_tokens.sh . grep -r "hf_" . --include="*.md" --include="*.py" || true hf upload beAnalytic/eda-llm-qwen2.5-lora . --exclude '.env' '*.env' ``` Substitua o nome do repo pelo valor em [`hf_registry.json`](../hf_registry.json) se diferente. Verificar: `hf repo list-files `. ### README do modelo (metadados YAML) O Hub mostra *YAML Metadata Warning: empty or missing yaml metadata* quando o `README.md` do repositório de modelo não tem [frontmatter YAML](https://huggingface.co/docs/hub/model-cards#model-card-metadata). Usa o template [`README_MODEL_CARD.md`](../README_MODEL_CARD.md): cola o conteúdo no editor do README no site ou, a partir da pasta de config, `hf upload README_MODEL_CARD.md README.md` (o terceiro argumento é o caminho no repo; ajusta org/repo). ## 5. Merge para modelo full (opcional) ```bash python scripts/promote_merged_model.py \ --base-model Qwen/Qwen2.5-1.5B-Instruct \ --lora-repo beAnalytic/eda-llm-qwen2.5-lora \ --merged-repo beAnalytic/eda-llm-qwen2.5-merged ``` ## 6. Git no clone do Space Trabalhe dentro do diretório do clone (ex.: `~/hf-spaces/Training`), não dentro de `ml/training/`. **Alterações não commitadas antes de pull:** ```bash git add -A git commit -m "Atualizar arquivos de treinamento" git pull --rebase origin main git push origin main ``` **Stash:** ```bash git stash git pull --rebase origin main git stash pop ``` **Branch divergiu:** ```bash git fetch origin git pull --rebase origin main git push origin main ``` **Rejeitou push:** ```bash git fetch origin git pull --rebase origin main git push origin main ``` Evite `git push --force` salvo decisão consciente de sobrescrever o remoto. ## 7. Setup inicial dos repos no Hub ```bash cd ml/configs/huggingface_training_config export HF_TOKEN=... python scripts/setup_huggingface.py ``` ## Referências - Métricas e TensorBoard: [METRICS.md](METRICS.md) - Problemas comuns: [TROUBLESHOOTING.md](TROUBLESHOOTING.md)