LMM-10K is a large-scale multimodal video dataset consisting of 10,000 high-quality 4K videos at 60 fps enriched with semantic, perceptual, and content-aware annotations.
Building a dataset of this scale required substantial data acquisition work. Leon Kordasch contributed the automation scripts that streamlined the entire acquisition pipeline. Reach out to Leon via email or GitHub below for questions about the pipeline or to see more of his work.
@inproceedings{ghasempour2026lmm10k,
title={{LMM-10K}: {Large}-{Scale} {4K} {Multimodal} {Dataset} for {Perceptual}, {Semantic}, and {Content}-{Aware} {Video} {Processing}},
author={Ghasempour, Mohammad and Wei, Yiying and Amirpour, Hadi and Timmerer, Christian},
booktitle={Submitted to ACM Multimedia 2026},
year={2026}
}
Mohammad Ghasempour
📧 mohammad.ghasempour@aau.at