black-forest-labs/flux-3-action-so101
Black Forest Labs released FLUX 3 Action SO-101, an open weights 7B world action model taking camera frames, robot state, and text instructions to predict actions. The repository provides loading instructions using LeRobot alongside recipes for fine-tuning, though applications must enforce hardware safety limits since the model outputs joint targets directly.
This system processes visual data from two cameras, robot state information, and text prompts to generate future motor commands. It simultaneously produces the next set of actions and predicted video frames using a denoising process. The architecture operates as part of a larger collection and relies on pre-saved processor pipelines for managing data normalization and observation contracts. Developers integrate the policy through a specific robotics library that handles the loading of frozen encoders and configuration. The provided recipe allows for fine-tuning the model on new datasets using a low-rank adaptation method. Execution involves predicting a specific number of actions, executing a portion of them at a fixed rate, and then replanning the remainder based on updated data. The publisher states that the model does not bound joint velocity, force, or workspace limits. Users must implement their own safety constraints and maintain a hardware stop mechanism to prevent dangerous movements. Validation through simulation or with safety limits engaged is required before operating near people. The license imposes restrictions on using the model for high-risk decisions or in ways that endanger individuals.
README
black-forest-labs/flux-3-action-so101 View on Hugging Face
Loading the README from Hugging Face…