XGENlabs/XGEN-JING

XGEN-JING releases an egocentric interactive experience model generating first-person video and audio from actions and reference images. Built on MiniMax-H3, it supports keyboard-controlled navigation and dialogue. The repository provides four-step bidirectional inference code, prompt skills, and validated setup instructions for six H100 GPUs. A causal model and technical report are pending, limiting current reproducibility.

README

XGENlabs/XGEN-JING View on Hugging Face

Loading the README from Hugging Face…