How Claude performs on robotics tasks

Anthropic researchers evaluate language model performance on robotics tasks using simulated and real robots. The study tests control methods ranging from direct motor torque commands to high-level policy supervision. Results indicate models succeed when supervising pretrained controllers but fail at direct joint control. The paper provides technical insights into model embodiment capabilities without releasing new model weights or developer tools.

Cover image for How Claude performs on robotics tasks