DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories

Chronological Source Flow
Back

AI Fusion Summary

Researchers introduced DeltaML-Bench, a new benchmark featuring 48 tasks from research papers to evaluate autonomous machine learning agents. These agents must navigate open-source repositories, repair pipelines, and improve baselines under compute constraints. Testing GPT-5 and Claude Sonnet 4 revealed that ARG scaffolding significantly improves performance; specifically, ARG increased the per-run success rate of GPT-5 from 9.4% to 33.9% within a 4 x 6h allocation. This benchmark addresses gaps in existing evaluation methods for ML experimentation.
Community Comments
Loading updates...
0