AMAP-ML/LoopArena

LoopArena provides a Python framework for benchmarking models as runtime controllers in loop engineering workflows. The repository documents evaluation methods for coding agents and software engineering tasks. It offers structured testing for agent behavior rather than a single model release. Developers can install the package to run standardized benchmarks on their own agent implementations.

README

AMAP-ML/LoopArena View on GitHub

Loading the README from GitHub…