QA Engineer
Storm3
Senior QA Lead (AI Systems)
About the Company
We’re a fast-growing Series A AI startup building software for the life sciences industry. Our platform helps enterprise customers create and review complex, regulated content using AI-powered workflows. We’re working with leading organizations in the space and are building at the forefront of LLMs, agentic systems, and applied AI.
About the Role
We’re hiring our first dedicated QA Lead to own quality across our AI platform. This is a hands-on engineering role focused on building testing infrastructure, evaluation systems, and automation that ensure both product reliability and AI output quality.
You’ll be responsible for everything from LLM evaluations and agent testing to Playwright automation, CI/CD quality gates, and release sign-off. The ideal candidate can write production-quality code, build scalable test systems, and think critically about how to measure quality in non-deterministic AI environments.
What You’ll Do
- Build evaluation frameworks for LLM and agent-based systems
- Design regression tests for model, prompt, and workflow changes
- Create automated testing infrastructure using Python and Playwright
- Own end-to-end testing across web applications and APIs
- Develop CI/CD quality gates and automated release workflows
- Perform manual and exploratory testing on new features
- Monitor production systems and convert failures into repeatable tests
- Test agent workflows for failure modes such as tool misuse, looping, hallucinations, and prompt injection
- Partner with engineering and product teams to establish quality standards
- Define and scale the QA function as the first dedicated hire
Required Experience
- Strong Python engineering skills
- Extensive Playwright and test automation experience
- Experience building and maintaining CI/CD testing pipelines
- Hands-on manual and exploratory testing expertise
- Familiarity with AI, ML, or LLM-based systems
- Experience designing evaluation methodologies and quality metrics
- Strong debugging, root-cause analysis, and error investigation skills
- Ability to operate independently and build processes from scratch
Preferred Experience
- Experience with LangChain, LangGraph, or agent frameworks
- Familiarity with LangSmith, Langfuse, Arize, Braintrust, Promptfoo, OpenAI Evals, or similar tools
- TypeScript and React experience
- Adversarial testing or red-team experience
- Backend experience with async Python services and job queues
- Experience in regulated industries such as healthcare, pharma, or finance
- MLOps or LLMOps experience
Compensation & Benefits
- $160,000-$300,000 base salary
- Meaningful equity package
- Comprehensive medical, dental, and vision coverage
- 401(k) with company match
- Visa sponsorship available
- Relocation support
- Unlimited PTO
- Professional development stipend
Location
New York City 5 days. Relocation assistance available
