Job Title: Software Engineer I (Research Engineer)
Location: Remote - California or PST time zone preferred
- Open to other time zones If stellar candidate, but has to be HIGHLY qualified.
Duration: 12 months, potential for extension.
Job Description:
Role Summary:
- We're looking for a software engineer to help build and evaluate CUA (Computer Use Agents) - models trained to operate real software and complete computer-based tasks the way a skilled engineer would.
- You'll own complex, long-running technical workflows end to end - implementation, validation, documentation, and follow-through - and design and maintain the tasks, benchmarks, and test suites that measure CUA capability as the underlying models, APIs, and infrastructure keep evolving.
- You hold an exacting bar for code quality: you love writing tests, you're fluent with deployment strategies like canary releases and rollbacks, and you bring a strong sense of risk management.
- You have a sharp eye for weak implementations and give (and take) blunt, direct feedback.
- You're also fluent with AI coding assistants like Claude Code, Codex, or Cursor, with your own well-developed best practices for using them, and you're proficient in Rust, Python, and/or TypeScript.
- Above all, you bring strong debugging and experimental discipline - the ability to investigate discrepancies across code, configuration, infrastructure, and results within the CUA stack, and turn ambiguous findings into reproducible conclusions.
Must-Have Skills
- Strong engineering background, AI model training.
- Python/torch
- Experience with ML evaluation, distributed systems, developer infrastructure, or large-scale testing.
- Experience maintaining benchmarks or test suites while underlying models, APIs, and infrastructure evolve.
Nice-to-have Skills:
- Rust and typescript
Years of Experience:
- 1-2 (3 years max) years of AI model training, or has extensive adademic background in engineering AI
Degrees/Certifications Required:
- Bachelors degree in a related field.
Mandatory Requirements:
- Passionate about coding since childhood with a drive to dive deep into technology. You maintain extremely high standards for code quality, love writing tests, and are familiar with deployment strategies like canary releases and rollbacks. You possess a strong sense of risk management.
- Familiar with at least one AI coding assistant (e.g., Claude Code, Codex or Cursor) and have developed your own insights and best practices for using them.
- Proficient in one or more of the following languages: Rust, Python, or TypeScript.
- Possess a sharp eye for spotting "garbage" code/implementation. You are willing to give blunt, direct feedback (call out bad code) and are equally thick-skinned enough to receive it.
- Strong debugging and experimental discipline. You can investigate discrepancies across code, configuration, infrastructure, and results, then turn ambiguous findings into reproducible conclusions.
- Comfortable owning complex, long-running technical workflows end to end, including implementation, validation, documentation, and follow-through.
Bonus Qualifications:
- Experience in server-side or infrastructure development, with a track record of building highly available and stable systems.
- Exceptional communication skills. Able to engage effectively with algorithm teams, product managers, operations, and executives. You can translate complex technical concepts into language that absolutely anyone (technical or non-technical) can easily understand.
- You have your own carefully maintained open-source project(s)-the number of GitHub stars doesn't matter.
- Experience with ML evaluation, distributed systems, developer infrastructure, or large-scale testing.
- Experience maintaining benchmarks or test suites while underlying models, APIs, and infrastructure evolve.
Key Projects/Day-to-Day Responsibilities:
- We're looking for a software engineer to help build and evaluate CUA (Computer Use Agents) - models trained to operate real software and complete computer-based tasks the way a skilled engineer would.
- You'll own complex, long-running technical workflows end to end - implementation, validation, documentation, and follow-through - and design and maintain the tasks, benchmarks, and test suites that measure CUA capability as the underlying models, APIs, and infrastructure keep evolving
Purpose/Size of this team & where does this position fit within the team?
- FAIR - training AI models, work with researchers investigation, enablement, creating benchmarks, evaluating model gaps.
Interview Process:
How many rounds of interviews? 2, zoom video interviews.
Types of Interviews: Behavioral and technical interviews
Interview Duration: 45 mins
|