Overview

LMM-10K is a large-scale multimodal video dataset consisting of 10,000 high-quality 4K videos at 60 fps enriched with semantic, perceptual, and content-aware annotations.

Dataset

⬇ Download Dataset (LMM-10K)

Acknowledgments

LK

Leon Kordasch

Building a dataset of this scale required substantial data acquisition work. Leon Kordasch contributed the automation scripts that streamlined the entire acquisition pipeline. Reach out to Leon via email or GitHub below for questions about the pipeline or to see more of his work.

Citation

@inproceedings{ghasempour2026lmm10k,
  title={{LMM-10K}: {Large}-{Scale} {4K} {Multimodal} {Dataset} for {Perceptual}, {Semantic}, and {Content}-{Aware} {Video} {Processing}},
  author={Ghasempour, Mohammad and Wei, Yiying and Amirpour, Hadi and Timmerer, Christian},
  booktitle={Submitted to ACM Multimedia 2026},
  year={2026}
}

Contact

Mohammad Ghasempour
📧 mohammad.ghasempour@aau.at