Credits
Most clips on Watch My Lips come from two research datasets whose makers share them openly. Thank you to the researchers, and to the speakers and actors who took part.
GRID audiovisual sentence corpus
The lip-shape, number and letter clips are single words cut from the GRID corpus: Martin Cooke, Jon Barker, Stuart Cunningham and Xu Shao, “An audio-visual corpus for speech perception and automatic speech recognition”, Journal of the Acoustical Society of America 120(5), 2421–2424 (2006). The corpus is licensed under Creative Commons Attribution 4.0 and available from the University of Sheffield and Zenodo.
Changes we made: each word is cut out of its sentence with a moment before and after, cropped to the speaker's mouth, resized and re-encoded.
CREMA-D
The everyday-sentence clips come from CREMA-D, the Crowd-sourced Emotional Multimodal Actors Dataset: Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova and Ragini Verma, “CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset”, IEEE Transactions on Affective Computing 5(4), 377–390 (2014). It is made available under the Open Database License, and its individual contents under the Database Contents License, from github.com/CheyneyComputerScience/CREMA-D.
Changes we made: we use the neutral take of each sentence, cropped to the actor's mouth, resized and re-encoded.
Software
Faces are found with YuNet (MIT licence) in OpenCV, and clips are converted with FFmpeg.
Questions
If you appear in a clip and want it removed, or you spot a credit we got wrong, email [email protected].