About the role
We are building the engineering substrate for agents that investigate unfamiliar systems, write and run code, create their own tools, and improve through experience. This role sits where emerging AI capabilities meet rigorous systems engineering. You will turn ideas that still feel like research into infrastructure people can understand, control and trust in production.
What you would do
- Design and build the runtime foundations that let autonomous agents act safely in real engineering environments.
- Solve hard problems at the boundary of model behaviour, deterministic software, isolation and developer experience.
- Turn experimental capabilities into simple, durable product primitives with explicit guarantees.
- Build feedback loops that make failures observable, reproducible and useful to both people and agents.
- Shape the architecture, interfaces and engineering culture of a platform at an early and consequential stage.
What we look for
- Tests you have seen go red by breaking the code on purpose.
- Respect for compatibility: a change that silently re-keys a cache or changes an output format is a regression.
- Clear explanations of why something is built the way it is.
- Restraint: no claim that cannot be backed by evidence.
Levels
| Level | What we expect |
|---|---|
| Junior | Reads unfamiliar code, explains it correctly, adds a test or a small fix. |
| Middle | Changes behaviour with tests, keeping existing guarantees. |
| Senior | Features and reviews that touch cache identity, limits or boundaries without regressions. |
| Staff+ | Designs isolation, failure handling and flows; knows what to refuse to promise. |
Example challenges
These are starting points, not a strict assignment list. Choose one that suits your interests and experience, adapt its scope, or propose a focused challenge of your own.
ENG-1 A usability review of zen
Easy Read zen --help and the help of three subcommands. Run each with a mistake: a typo, a
missing argument, a flag that does not exist.
Deliver five specific problems (unclear wording, a misleading error, a missing hint), each with the exact text and a proposed replacement.
ENG-2 Read the sandbox status
Easy Run zen sandbox status in a scaffolded project. Explain, line by line, what it reports
about the connection, the capacity and the container, and what each number would change if you
raised or lowered it.
Deliver the output and your explanation, with one thing in it you could not explain.
ENG-3 Fix a papercut
Medium Find one small flaw in a command's output, error message or help, fix it, and add a test
in packages/.
Deliver the change as a pull request: the before and after output, the test, and why the fix does not break anyone who parses the old output.
ENG-4 Follow one request
Medium Start the mock, send one request to an operation that has never been called, and trace it.
Deliver a short written trace: where the generator is written, where it is judged, where it is cached and under what key, which process runs it, and what the second identical request skips. Name the files. Then answer: why does the serving container have no network, and why is it not persisted across starts?
ENG-5 Pin a behaviour with a test
Medium The sandbox runner retries commands that fail for transport reasons (an ssh handshake
dying, a connection reset). It must never repeat a command that already produced output. Read the
code, then write tests that prove both halves with an injected runner (see
packages/). Break the guard on purpose and show your test goes red.
Deliver the tests, and the mutation you used to check them.
ENG-6 Try to escape the file tools
Medium The file tools are confined to a workspace. Write adversarial tests: .. segments,
symlinks (existing and created mid-run), absolute paths, the mount prefix, odd spellings of the root,
and moves and copies whose source is read-only.
Deliver the tests, the result of each, and a report on any escape. If there is none, say what your tests could not reach.
ENG-7 Make the mock produce a very long list
Hard Today a paged list ends after a handful of pages, which is right for most clients and useless for testing a client's behaviour on a thousand. Add a way to make one operation serve a chosen number of pages (for example 200), with the end reached cleanly and the token still advancing and never repeating.
- The cache key of every operation that is not affected must stay byte-identical. There are
pinned hashes in
packages/; do not re-pin them to make a test pass.faker/ test/ paging.test.ts - The server has a request-time backstop against looping tokens. Yours must pass through it.
- The generator runs without a network and without remembering anything between requests.
Deliver the change, tests, and a note on where the page limit is decided today and why you changed it there and not somewhere else.
ENG-8 Wide characters in the terminal
Hard The interface keeps its frame within the window height by measuring text and wrapping it
(see packages/). Check what happens when the text holds double-width characters
(CJK, emoji) or combining marks. A frame taller than the window leaves stray copies of its top lines
on screen.
Deliver a test that fails today (or a demonstration that it does not, and why), the fix, and what you chose not to handle.
ENG-9 A hostile archive
Hard zen import unpacks a project archive. Build hostile archives: path traversal, symlink
entries, absurd entry counts, an expansion bomb, names that differ only in case, odd permissions.
Deliver the archives (generated by a script, not committed as binaries), what the import did with
each, and a report on any that got through. Existing tests in packages/
are a starting point, not a limit.
ENG-10 Egress
Expert The strict hardening profile defaults the network to none, and an egress allow-list
is explicitly out of scope. A customer wants a sandbox that can reach exactly one internal host and
nothing else.
Deliver a design, not code: threat model, where enforcement could live (image, container, host, proxy) and which of those a rootless podman machine on macOS can actually do, what the agent sees when something is blocked, how you would test the block, and what you would refuse to promise.
ENG-11 The first five minutes
Expert Someone installs the CLI for the first time with no key, no podman, and a terminal that is not a TTY (a CI job). Design the path from install to a first successful run.
Deliver a design: what is detected, what is prompted, what is skipped in CI, what each failure says and how it recovers, and how you would test the flow without a real machine for every case.