UniMate Animates Arbitrary Skeletons From Text With No Per-Skeleton Fine-Tuning
arXiv 2609.05415 addresses the gap left by automatic rigging: animation-ready 3D assets now exist at scale, but learned animators stay topology-constrained, needing category-specific templates or per-skeleton fine-tuning plus reference motions at inference. UniMate is a topology-aware diffusion transformer that synthesizes articulated motion for arbitrary skeletons from a rigged asset and a text prompt with no test-time optimization. It injects skeletal topology into attention three ways: a graph-aware attention bias from pairwise joint relations and geodesic distances, a spectral rotary position embedding generalizing RoPE to arbitrary kinematic trees via the graph Laplacian, and a global topological conditioner. It carried 13 upvotes on HuggingFace Daily Papers.
↳ Follow the thread