Bespoke Labs is seeking a researcher to develop and evaluate reinforcement learning environments and benchmarks for long-horizon agent tasks. The role requires demonstrated experience with multi-step reasoning agents, specifically mentioning contributions to environments like SWE-bench and Gymnasium. The position is a remote contract, with applications filtered based on concrete evidence of long-horizon agent work. AI
IMPACT This role focuses on advancing agent capabilities for complex, long-term tasks, potentially leading to more sophisticated AI systems.
RANK_REASON Job posting for a specific role related to AI agent development and evaluation.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →