A 2025 self-improving-agent experiment raised a coding benchmark score from 14.2% to 30.7%, but this is not evidence of unlimited autonomous…