AI Models Tested on Recursive Self-Improvement Benchmarks

Researchers have developed a new benchmark, HarnessOpt-Bench, to evaluate the ability of advanced AI models to improve the performance of other AI agents without access to test data. This approach aims to measure recursive self-improvement in a controlled environment, preventing the AI from simply memorizing solutions.

The benchmark isolates the AI optimizer from sensitive information, ensuring that improvements are genuine and not a result of cheating. The system uses external servers for final scoring and enforces budget limitations, creating a secure sandbox for the evaluation process. This method is designed to provide a more accurate assessment of an AI's capacity for genuine self-enhancement.

Initial tests on five frontier large language models (LLMs) revealed varying degrees of success. Claude Opus 5, when paired with OpenCode, outperformed other models on three out of four tasks. The study also tracked the progress of GPT models on a specific task over several months, showing a significant increase in performance from 3% to 49% of the head, indicating potential for substantial AI-driven improvements.

28 stories · 7 sources

#deepseek #ai #benchmarks

Other digests