← Back to home
Careers
Build the hardest tasks in AI.
We are building a small, vetted team of engineers who can design coding and agent tasks that strong AI models mostly fail, and prove it with automatic checkers. If you write clean code, think hard about correctness, and enjoy breaking good models, we want to talk.
What you would do
- Design self-contained coding, RL-environment, RLVR, or agent tasks on real open-source code.
- Write deterministic tests and checkers that fail on the broken code and pass only on the fix.
- Tune difficulty against strong agents and find the root cause of every failure.
- Hold the bar. Every task you ship clears our checks before a client sees it.
What we look for
- Strong engineering in Python, JS or TS, Go, or Rust, and comfort with Docker.
- A real feel for test design and edge cases. You can make a task hard and still fair.
- Clear writing. Specs that say what to do, not how to do it.
- A plus: past work on public coding or agent benchmarks.
How the team works
We run a short, paid trial task, scored on the same bar we hold ourselves to. Strong work earns a place on the team, steady assignments, and review responsibility as you build a track record. It is remote and flexible. We care about what clears the checks, not hours logged.
Apply, tell us what you have broken
Send a short note and links to your best work to akash@topdeta.com.