Completed Computing & AI Bones, Joints & Muscles

DEFORM: Large Scale Shape Analysis of Deformable Models of Humans

In plain English

AI plain-English summary

Computer vision systems currently recognise people in photos by spotting sparse landmarks—eyes, elbows, joints—but they cannot grasp the full 3D shape of a moving body. This project will build the first large-scale database of high-resolution 4D scans (3D over time) of humans, then develop algorithms to map those dense 3D shapes onto ordinary 2D images captured “in the wild.” The core problem is a data gap: collecting thousands of detailed 3D scans of people in motion remains expensive and slow, so existing systems rely on cheap 2D images annotated with only a few dozen points. Without dense 3D information, machines cannot truly understand how a person’s body deforms during a gesture, a fall, or a dance. If successful, DEFORM would let a computer watch a single 2D video and reconstruct the full 3D shape and movement of every person in it. That could transform medical gait analysis, sports coaching, virtual try-on for clothing, and human-robot interaction—any system that needs to interpret what a body is doing, not just where its joints are. The work is primarily fundamental science, building the geometric foundation for a generation of vision systems that see people as they actually are.

View original technical description
Recently, computer vision is witnessing a paradigm shift. Standard robust features, such as Scale Invariant Feature Transform (SIFT), Histogram of Oriented Gradienst (HoGs), etc., are replaced by learnable filters via the application of Deep Convolutional Neural Networks (DCNNs). Furthermore, for applications (e.g., detection, tracking, recognition, etc.) that involve deformable objects, such as human bodies/faces/hands etc., traditional statistical or physics-based deformable models are combined with DCNNs with very good results. The current progress is made due to the abundance of complex visual data in the Big Data era, spread mostly through the Internet via web services such as Youtube, Flickr, and Google Images. The latter has led to the development of huge databases (such as ImageNet, Microsoft COCO, and 300W, etc.) consisting of visual data captured "in-the wild". Furthermore, the scientific and industrial community has undertaken large-scale annotation tasks. For example, me and my group have made huge efforts to annotate over 30K facial images and 500K video frames with regards to a large number of facial landmarks. The COCO team has annotated thousands of body images with regards to body joints, etc. All the above annotations generally refer to a set of sparse parts of objects and/or their segments, which can be annotated by humans (e.g., through crowd sourcing). In order to make the next step in automatic understanding of a scene in general, and humans and their actions, in particular, the community needs to acquire 3D dense information. Even though the collection of 2D intensity images is now a relatively easy and inexpensive process, the collection of high-resolution 3D scans of deformable objects, such as humans and their (body) parts, still remains an expensive and laborious process. This is the principal reason why very limited efforts have been made in collecting large-scale databases of 3D faces, heads, hands, bodies, etc. In DEFORM, I propose to perform large-scale collection of high-resolution 4D sequences of humans. Furthermore, I propose new lines of research in order to provide high quality annotations regarding the correspondences between the 2D intensity "in-the-wild" images and the dense 3D structure of deformable objects' shapes and in particular of humans and their parts. Establishing dense 2D-to-3D correspondences can effortlessly solve many image-level tasks such as landmark (part) localisation, dense semantic part segmentation, estimation of deformations (i.e., behaviour), etc.

View the original record at the funder ↗

Researchers

Stefanos Zafeiriou (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

Learning compact and efficient deformable models of human shape variation
A 4D Foundation for Deformable Objects
3D Shape Understanding using Deep Learning
3D Intrinsic Shape Recognition Under Deformation and View Changes
High Quality 3D Geometry and Appearance Reconstruction of Non-Rigidly Deforming Objects using Low-Cost RGB-D Cameras

Original classification

Fellowship

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.