Researchers have developed a self-supervised framework called software-in-the-loop reconstruction (SWR) to train terminal agents for scientific domains. This method leverages existing scientific software workflows to generate reference outputs and verification targets, reducing the need for manual engineering. By executing multiple input configurations and partitioning cases, SWR enables agents to construct editable programs without direct access to source code, which are then evaluated against workflow outputs. The framework has been instantiated with 500 workflows across six domains, and fine-tuning a Qwen3.8-27B model using SWR-generated data improved its performance on the Terminal-Bench benchmark. AI
IMPACT This approach could enable more efficient training of AI agents for specialized scientific tasks by leveraging existing codebases.
RANK_REASON The cluster contains an academic paper detailing a new framework and experimental results. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX Code Finder for Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Qwen3.8-27B
- Qwen3.8-Max
- ScienceCast
- software engineering
- Terminal-Bench
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →