About the role
Testing intelligent systems is still an open engineering problem. Correctness is no longer only a known input and a fixed output: an agent can take a novel path, produce a plausible answer and still cross an important boundary. You will help discover how to measure intelligence within constraints, how to test the limits of evaluation harnesses, and how to build evidence that an adaptive system is not only capable, but dependable.
What you would do
- Invent evaluation strategies for behaviour that is generated, non-deterministic and context dependent.
- Probe the boundaries between models, agents, tools, harnesses and real systems to discover where confidence breaks down.
- Develop useful oracles, invariants and adversarial scenarios when there is no single expected answer.
- Turn subtle failure modes into clear evidence and automated checks that teams can trust.
- Help define what quality means for agentic software, including which guarantees matter and which would create false confidence.
What we look for
- Reproducible reports: steps, expected, actual, evidence.
- Tests that fail loudly and say why, and suspicion of green results: a test that cannot fail proves nothing.
- Judgement about severity, and the nerve to say "this is not a bug".
Levels
| Level | What we expect |
|---|---|
| Junior | Systematic exploration, clear reports, readable repeatable tests, checking before claiming. |
| Middle | Reports that classify cause (contract, semantic, state); helpers and fixtures others reuse. |
| Senior | A test strategy with a defended cut-off; oracles for non-deterministic output; CI that stays trusted. |
| Staff+ | Sets quality criteria and a test architecture for a product whose behaviour is partly generated. |
Example challenges
These are starting points, not a strict assignment list. Choose one that suits your interests and experience, adapt its scope, or propose a focused challenge of your own.
QA-1 Smoke and map
Easy Call every operation in mail.yaml and calendar.yaml at least once with valid input. Include at least one call that should fail (missing required body field, unknown path, wrong method) and say what status you expected before you made it.
Deliver a table: operation, request, status, one thing that surprised you.
QA-2 Check the documentation's claims
Easy Read packages/faker/README.md. Pick ten claims about observable behaviour (headers, status codes, endpoints, how paging ends). Check each against the running mock.
Deliver a table: claim, request you used, result, verdict (true, false, cannot check).
QA-3 A first test that guards the routes
Easy Write a test that starts from the two specs and asserts that GET / lists
every operationId in them, and that GET / answers. It must fail with a readable
message if one is removed from the spec.
Deliver the test and the run output, passing and deliberately failing.
QA-4 Walk a list to its end
Medium GET / is paged. Follow nextPageToken until it stops.
Record, per page, the item count, the token, and whether any id was already seen on an earlier
page. Do the same for GET / and say which behaves better. Find one list
in the specs that is not paged and show that the mock does not pretend otherwise.
Deliver the walks, the total pages, and a verdict on resultSizeEstimate: does it match what
you counted?
QA-5 Exploratory test of the contents page
Medium GET / lists every operation with a Run button and a form per parameter. Test the
form as a user would: required fields marked, prefilled bodies accepted, errors readable, enum and
number bounds respected, bad input handled.
Deliver a list of defects and papercuts, each with steps, expected, actual, severity and a screenshot or copied response.
QA-6 Contract suite
Medium Write a suite that calls every operation in both specs and validates each response body against that operation's response schema taken from the spec itself (do not hand-copy schemas).
Deliver the suite and a run report. It must fail loudly, naming the operation and the violated keyword, when you break a response on purpose.
QA-7 A pager that cannot hang
Medium Write walk(url, opts) that follows nextPageToken to the end and returns every item.
It must stop on its own when the server loops (same token twice), when a token was seen before,
and after a hard page cap, and it must say which of those happened. Run it against listMessages,
listEvents and listThreads.
Deliver the helper, tests for its three stop reasons using a stub server of your own, and the
result against the mock. Include a negative case: listLabels is not paged.
QA-8 Does the search find anything?
Hard Listing returns ids only, so judging a search takes two calls: list, then getMessage
per id. Run these and judge the returned messages, not the status code:
q=from:ada@example.orgq=is:unreadq=after:2026/01/ 01 before:2026/ 02/ 01 - calendar
q=standup, andtimeMin/timeMaxaround one week
Deliver a defect report for each query that fails you: steps, expected, actual, severity, and a one-line argument for whether the mock should be able to honour it. Be ready to be asked what the spec actually promises.
QA-9 Determinism, flake and cold starts
Hard Start the mock with --seed 42 and prove, with numbers, that the same request gives the
same body: twenty times in a row, after a restart, and fifty at once. Then do it cold (empty
cache: zen faker cache clear) and explain what changed. The first call to an operation can cost a
model turn.
Deliver the tests, the cold-start findings, and a rule for your team: what a test may assume
about timing and cache headers (x-faker-cache), and what it must never assume.
QA-10 Run it in CI with no model
Hard ci.yml runs without provider keys. A suite that needs a live mock needs generators, which a model writes. Make a suite from QA-6 to QA-9 run in CI.
Deliver the approach (for example pre-written generators checked in as fixtures, or a stub in front of the model), the change to the workflow, and an honest statement of what the CI run proves and what it cannot.
QA-11 What should a mock promise?
Expert Nobody has told you what "adequate" means for generated data. Decide.
Deliver one page:
- Promises a mock could make, grouped as contract, echo (a path value comes back), semantic (a search result matches its query), referential (an id from a list resolves), and stateful (a sent message appears in the list).
- For each: how you would test it, what it would cost the mock to keep it, and whether you would require it.
- Three acceptance criteria written so that two people would always agree on pass or fail.
No right answer exists for the cut-off. Defend yours.
QA-12 Oracles for data that has no fixed answer
Expert You cannot assert the exact body of a generated search. You can assert relations, for example:
maxResults=nreturns at mostnitems;- adding a clause to
qnever grows the result set; - the union of all pages of a query equals the first page of the same query with no cap;
- for calendar, every returned event overlaps
[timeMin, timeMax); freeBusy.calendarscontains exactly the ids you asked about.
Run them and report a pass rate per relation.
Deliver the suite plus a short note: for each failing relation, is it a mock defect, a test defect, or a thing a stateless mock cannot do? What goes into quarantine and what blocks a release? Show one test you removed because it was wrong.