An Anthropic researcher just gave us a peek at self-improving AI
Anthropic researchers published a study on automated alignment post-training. The system iteratively proposes and tests methods to fix misalignment benchmarks. It outperformed human-led efforts in speed and cost. No usable model weights or software artifacts are released for developers to integrate into their own pipelines.
