Abstract

Current systems configured to generate three-dimensional interactions between a hand and an object typically rely on fragmented computational pipelines that can fail to generalize to novel target objects and may offer restricted control. A system is disclosed for generating poses of a hand utilizing a hand diffusion transformer configured to process conditioning signals, such as target object geometry and a direction of a hand. The hand diffusion transformer may output a distance profile, which represents target distances between regions of a hand and surfaces of an object. A subsequent optimization layer can use the distance profile to refine candidate hand poses to match the target distances. By predicting a set of distances rather than direct coordinates, the system can facilitate desired target contacts while mitigating collisions. The disclosed system may provide controllable hand pose generation configured to adapt to objects having varying dimensions and geometries.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS