About the role
Agentic software changes what delivery infrastructure must guarantee. Systems that write code, create tools and evolve their own behaviour need an engineering foundation that can move quickly without trading away trust. You will shape that foundation: the path from an idea to a verifiable, secure and repeatable release of technology that is defining a new category.
What you would do
- Build delivery systems that let a fast-moving agentic AI platform evolve with confidence.
- Design resilient release flows that remain understandable and recoverable when the unexpected happens.
- Turn reproducibility, observability and supply-chain integrity into everyday engineering advantages.
- Find the deeper causes of unreliable tests and environments, then eliminate whole classes of failure.
- Shape the standards and automation that help every engineer ship ambitious work safely.
What we look for
- Evidence over opinion: timings, run logs, counts.
- Changes that are safe to re-run and say what they did.
- Respect for the people who debug your pipeline at a bad moment.
Levels
| Level | What we expect |
|---|---|
| Junior | Reads the pipeline, explains it correctly, spots problems. |
| Middle | Improves a workflow with proof; fixes a flaky test at its cause. |
| Senior | Designs for partial failure; hardens supply-chain and reproducibility. |
| Staff+ | Owns release and quality infrastructure for the whole product. |
Example challenges
These are starting points, not a strict assignment list. Choose one that suits your interests and experience, adapt its scope, or propose a focused challenge of your own.
OPS-1 Draw the release
Easy Read the release script and workflow. Draw the path from npm run release to a published
package as a Mermaid diagram (npm run mermcheck validates Mermaid in this repository).
Deliver the diagram, and three ways the process could fail, with what you would see in each case.
OPS-2 Tidy a workflow
Medium Audit the workflows for repeated work, missing caching, slow steps and steps that cannot fail. Change one thing.
Deliver the change, the run times before and after (from real runs), and what you left alone.
OPS-3 Find the flaky tests
Hard Run the offline suite (npm test -- --exclude '**/) many times, and also in a clean
home directory and with a different timezone. Record which tests fail and how often.
Deliver a table of failures with causes (timing budgets, dependence on the home directory, shared state, order), and a fix for one at its cause rather than by raising a timeout.
OPS-4 Pin the sandbox image
Hard The sandbox image builds from a Dockerfile that installs packages at build time. Make it reproducible: pinned base image, pinned dependencies, a record of what went in.
Deliver the change, how you check that two builds match, and what you cannot pin and what that costs.
OPS-5 A release that can fail half-way
Expert Two packages publish, the third fails. Review the current process and design a release that is safe to re-run, cannot publish two tags at once, and has a clear rule for taking a bad version back.
Deliver a design with the failure cases you considered, what is automatic and what needs a person, and the one change you would make first.