MC3-18 HMDB51 (UCF-101 Init)
MC3-18 (Mixed Convolution 3D) fine-tuned on HMDB51 split 1, initialized from this project's own MC3-18/UCF-101 model (87.05% accuracy) instead of Kinetics-400 -- trained as part of the video pipeline in human-action-classification, to test whether a domain-closer pretraining source (another trimmed, YouTube-sourced action dataset) transfers better than a larger but more generic one. A sibling model initialized from Kinetics-400 is also available; see Related Resources below.
Performance
| Metric | Value |
|---|---|
| Accuracy (Top-1) | 55.46% |
| Precision (macro) | 53.89% |
| Recall (macro) | 55.44% |
| F1 Score (macro) | 53.66% |
| Parameters | 11.5M |
| Best epoch | 49 / 100 |
Within 1 point of the Kinetics-400-initialized sibling model (56.34% accuracy) despite UCF-101 being a ~30x smaller pretraining corpus -- see Kinetics-400 vs. UCF-101 Initialization below.
Evaluation Protocol
Metrics above come from VideoTrainer.validate() in hac.video.training.train, run on HMDB51 split 1's test set (1,530 videos, 51 classes), at the checkpoint's best-performing epoch. Each clip: 16 frames sampled at stride 2 (i.e. spanning up to 32 source frames), resized preserving aspect ratio to roughly 128x171, center-cropped to 112x112, normalized with Kinetics-400 statistics -- a single center clip per video, no test-time augmentation or multi-crop averaging.
Usage
Install Dependencies
Not yet published on PyPI -- install from source:
git clone https://github.com/dronefreak/human-action-classification
cd human-action-classification
pip install -e .
Load the Model from Hugging Face
import json
import torch
from huggingface_hub import hf_hub_download
from hac.video.models.classifier import Video3DCNN
config_path = hf_hub_download(repo_id="dronefreak/mc3-18-hmdb51-ucf-transfer", filename="config.json")
weights_path = hf_hub_download(
repo_id="dronefreak/mc3-18-hmdb51-ucf-transfer",
filename="mc3-18-hmdb51-ucf-transfer.pth",
)
with open(config_path) as f:
config = json.load(f)
model = Video3DCNN(
num_classes=config["num_classes"], # 51
model_name=config["model_type"],
pretrained=False,
)
checkpoint = torch.load(weights_path, map_location="cpu", weights_only=False)
model.load_state_dict(checkpoint["model_state_dict"])
model.eval()
Run Inference on a Video
The repo's VideoPredictor wraps frame sampling, transforms, and the forward pass end-to-end (pass num_frames=16 to match this model's training configuration):
from hac.video.inference.predictor import VideoPredictor
predictor = VideoPredictor(model_path=weights_path, num_frames=16, device="cpu")
result = predictor.predict_video("path/to/video.mp4", top_k=5)
print(result["top_class"], result["top_confidence"])
Note: VideoPredictor's built-in class list defaults to UCF-101's 101 classes -- for HMDB51 you'll want to pass/override the 51 class names listed below rather than relying on the predictor's default.
Training Configuration
| Setting | Value | Source |
|---|---|---|
| Dataset | HMDB51 split 1 (3,570 train / 1,530 test videos, 51 classes) | HMDB51 split files |
| Architecture | MC3-18 (torchvision.models.video.mc3_18) |
checkpoint config |
| Pretrained init | This project's MC3-18/UCF-101 model | checkpoint config + repo history |
| Optimizer | SGD (momentum=0.9, nesterov=False) | checkpoint optimizer state |
| Initial learning rate | 0.0003 | checkpoint optimizer state |
| Weight decay | 0.002 | checkpoint optimizer state |
| LR schedule | StepLR (step_size=20, gamma=0.1) | checkpoint scheduler state |
| Epochs trained | 100 (best at epoch 49) | checkpoint + training history |
| Frames per clip | 16, frame_interval=2 | training script default |
| Spatial resolution | 112x112 (aspect-preserving resize + random crop) | training script default |
| Batch size | not recorded in checkpoint | -- |
| Augmentation | MixUp (alpha=0.4), CutMix (alpha=0.8), label smoothing (0.1), RandomHorizontalFlip, ColorJitter, RandomGrayscale | training script default (unconfirmed exact values for this run) |
Rows marked "checkpoint ..." are read directly out of the optimizer/scheduler state and config dict stored inside mc3-18-hmdb51-ucf-transfer.pth. Rows marked "training script default" reflect hac.video.training.train's CLI defaults/flags at the time of training but weren't independently re-derived from the checkpoint for this exact run -- no separate run-config file was saved alongside it.
Kinetics-400 vs. UCF-101 Initialization
This project also ships an MC3-18/HMDB51 model initialized from Kinetics-400 instead of UCF-101 -- see mc3-18-hmdb51-kinetics (56.34% accuracy).
| Initialization | Accuracy | Notes |
|---|---|---|
| UCF-101 (this model) | 55.46% | ~30x smaller pretraining corpus than Kinetics-400; domain-closer to HMDB51 (similar YouTube/movie sources, overlapping action categories); 16-frame clips |
| Kinetics-400 | 56.34% | Larger, more diverse pretraining corpus; 8-frame clips (avoids tiling on HMDB51's shorter videos) |
The two reach nearly identical validation accuracy despite very different pretraining sources -- consistent with the idea that domain similarity can partly substitute for pretraining-set size, though a single run per initialization isn't enough to call that conclusive. The original training run's console logs reportedly showed a smaller train/validation gap for this UCF-101-initialized model than for the Kinetics-initialized one; that figure isn't stored in the checkpoint itself, so it isn't independently re-verified in this card.
Frame tiling caveat: this model uses 16-frame clips at stride 2 to match its UCF-101 pretraining configuration, but many HMDB51 videos are shorter than the resulting 32-frame span -- short clips get frame-repeated ("tiled") to reach 16 sampled frames, which may hurt performance on those specific samples. The Kinetics-initialized sibling avoids this by using 8-frame, stride-1 clips instead.
HMDB51 Classes
The model predicts 51 action classes: brush_hair, cartwheel, catch, chew, clap, climb, climb_stairs, dive, draw_sword, dribble, drink, eat, fall_floor, fencing, flic_flac, golf, handstand, hit, hug, jump, kick, kick_ball, kiss, laugh, pick, pour, pullup, punch, push, pushup, ride_bike, ride_horse, run, shake_hands, shoot_ball, shoot_bow, shoot_gun, sit, situp, smile, smoke, somersault, stand, swing_baseball, sword, sword_exercise, talk, throw, turn, walk, wave.
Known Limitations
- Frame tiling on short HMDB51 clips (see caveat above) may depress accuracy on a subset of test videos.
- Single model, no ensembling; no test-time augmentation (multi-crop, multi-clip temporal sampling).
- Trained and evaluated on HMDB51 split 1 only -- performance on splits 2/3 is unverified.
- Depends on this project's own MC3-18/UCF-101 checkpoint as its pretraining source rather than a widely-used public pretrained model, making external reproduction harder without first training that upstream model.
Repository Contents
mc3-18-hmdb51-ucf-transfer.pth
config.json
README.md
config.json doubles as the Hub's download-count query file: since this repo has no library_name integration the Hub recognizes, it falls back to counting requests against config.json (per Hugging Face's download-stats docs) -- the loading snippet above fetches it as part of normal usage, so downloads register.
Related Resources
- mc3-18-hmdb51-kinetics -- sibling model, same architecture/dataset, initialized from Kinetics-400 instead of UCF-101
- mc3-18-ucf101 -- the UCF-101 model this model was transferred from
- human-action-classification -- the training/inference framework used to produce this checkpoint
Citation
If you use this model, please consider citing the HMDB51 and UCF-101 datasets, the MC3 architecture, and the training framework:
@inproceedings{kuehne2011hmdb,
title={HMDB: a large video database for human motion recognition},
author={Kuehne, Hildegard and Jhuang, Hueihan and Garrote, Est{\'\i}baliz and Poggio, Tomaso and Serre, Thomas},
booktitle={2011 International Conference on Computer Vision},
pages={2556--2563},
year={2011},
organization={IEEE}
}
@article{soomro2012ucf101,
title={UCF101: A Dataset of 101 Human Actions Classes From Videos in the Wild},
author={Soomro, Khurram and Zamir, Amir Roshan and Shah, Mubarak},
journal={arXiv preprint arXiv:1212.0402},
year={2012}
}
@inproceedings{tran2018closer,
title={A Closer Look at Spatiotemporal Convolutions for Action Recognition},
author={Tran, Du and Wang, Heng and Torresani, Lorenzo and Ray, Jamie and LeCun, Yann and Paluri, Manohar},
booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2018}
}
@misc{saksena2025mc3hmdbucf,
author = {Saumya Saksena},
title = {{MC3-18 HMDB51 (UCF-101 Init)}},
year = {2025},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/dronefreak/mc3-18-hmdb51-ucf-transfer}},
note = {Trained with the human-action-classification framework, Top-1 Accuracy: 55.46\%}
}
@software{saksena2026hac,
author = {Saumya Saksena},
title = {{Human Action Classification: Pose-based and Video-based Models}},
year = 2026,
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/dronefreak/human-action-classification}}
}
License
Apache-2.0
- Downloads last month
- 19
Papers for dronefreak/mc3-18-hmdb51-ucf-transfer
UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Evaluation results
- Top-1 Accuracy on HMDB51test set self-reported55.460
- Macro Precision on HMDB51test set self-reported53.890
- Macro Recall on HMDB51test set self-reported55.440
- Macro F1 on HMDB51test set self-reported53.660