Quality
AI Evaluation Engineer
Write the cases that catch fluent lies. Gold sets, red teams, and a gate that can fail the build.
2 weeks · gold set, red team, a gate
0/11 lessons on this path
US
$110K–$230K+
Canada
CA$95K+
India
₹12–32 LPA+
What this path is (and is not)
Enough to catch fluent lies before a demo. Mastery is the set that still works next quarter.
Mastery syllabus
1 · Gold cases
Easy, hard, empty, fluent lie.
2 · The gate
Fail the build. Show the miss.
3 · Gold and judge
Twenty cases. A judge is a sieve, not a court.
4 · Catch the lie
The invented citation goes first in the file.
Simulation labs
EvalSuite
Twenty cases. One must refuse.
Open the labGround or refuse
No chunk, no decoration.
Open the labTwenty cases
Five of four kinds.
Open the labRater guide
Two humans, one lie.
Open the lab
The project that hires
Twenty cases and a gate that can fail the demo
- Easy, hard, empty, lie — five each.
- A rater guide of one page.
- One case that caught a real invented citation.
Interview room
- Write one fluent-lie trap for a handbook bot.
- How do you score groundedness with a human rater guide?
- Empty retrieval: what must the system do, and how do you test it?
- LLM-as-judge disagrees with two humans. Who wins?
Resume lines that hire
- Wrote a 20-case eval that caught invented citations before a client demo.