Software Development Engineer Test
Job Summary
At Penguin Solutions (Nasdaq: PENG) The AI Factory Platform Company were building a team of innovators who thrive on collaboration creativity and the opportunity to help shape the future of AI. As part of the AI technology revolution our teams design build deploy and manage AI factories for enterprises sovereign AI initiatives and neocloud providers worldwide.
Headquartered in Silicon Valley California Penguin Solutions operates globally through a network of R&D manufacturing and sales locations. For nearly three decades we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads from training to inference and agentic AI at scale.
Penguin Solutions brings together differentiated infrastructure software advanced memory compute systems end-to-end services and industry-leading partner solutions in a full-stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision.
At Penguin Solutions we value ideas over hierarchy and believe in servant leadership where leaders enable teams to do their best work. We empower employees to take ownership drive innovation and grow through challenging work continuous learning and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes Penguin Solutions is a place to do your best work grow your career and make a meaningful impact.
Penguin Computing is seeking an experienced AI QA / SDET Engineer to join our Software Engineering team. Penguin Computings Scyld Software products are used to deploy provision manage and monitor some of the largest computational systems in the world. As we expand these platforms with AI/LLM-powered and agentic capabilities this role will drive quality engineering across both our core infrastructure software and emerging AI solutions.
You will join our agile Software team and collaborate closely with Software Engineers AI Engineers Product Owners Product Managers and other teams across the organization to ensure our software meets the highest standards of quality reliability performance and security.
The ideal candidate combines a strong foundation in software testing and automation with hands-on knowledge of LLM evaluation AI agent testing AI security testing and modern AI quality methodologies. You will help define testing strategies build scalable automation and AI evaluation frameworks and continuously improve our testing processes and tools.
This role requires strong technical problem-solving skills initiative and the ability to work with a high degree of independence. You will have the opportunity to shape our evolving AI quality engineering practice and ensure that both traditional software and AI-powered capabilities are reliable secure measurable and production-ready.
- Define and execute end-to-end QA strategy for software AI/LLM features and AI agents.
- Design and maintain scalable test automation frameworks covering API UI integration regression performance and infrastructure testing.
- Build LLM/AI evaluation frameworks and automated evaluation pipelines integrated with CI/CD.
- Define evaluation datasets test scenarios golden datasets and quality benchmarks for AI features.
- Evaluate LLM outputs for correctness relevance groundedness hallucination consistency and task completion.
- Implement automated evaluation methodologies including LLM-as-a-Judge deterministic checks model-based grading and human evaluation.
- Test AI agents and agentic workflows including reasoning tool/function calling multi-step execution state management error handling and recovery.
- Validate RAG pipelines including retrieval quality context relevance grounding citation accuracy and response quality.
- Perform adversarial and security testing of AI systems including prompt injection jailbreaks data leakage unsafe tool execution excessive agency and access-control violations.
- Apply AI security methodologies such as OWASP Top 10 for LLM Applications and use appropriate AI red-teaming/security testing tools.
- Build automated regression suites to identify behavioral and quality degradation across prompt model agent and application changes.
- Define AI quality metrics dashboards release gates and acceptance criteria.
- Test AI systems across different models and configurations to validate reliability and model-independent behavior.
- Replicate customer environments through simulated data and usage patterns for realistic end-to-end testing.
- Partner with Software Engineering AI Engineering Product and Architecture teams to embed quality engineering early in the SDLC.
- Troubleshoot complex issues across Linux Kubernetes networking infrastructure and AI application layers.
Experience with one or more of the following is highly desirable:
- LLM Evaluation: LLM-as-a-Judge golden datasets pairwise evaluation rubric-based scoring human-in-the-loop evaluation
- Evaluation Frameworks: DeepEval Ragas Promptfoo LangSmith OpenAI Evals or equivalent
- Agent Testing: tool/function-call validation trajectory evaluation task-completion testing multi-agent workflow testing
- RAG Evaluation: retrieval precision/recall context relevance faithfulness/groundedness answer correctness
- AI Security: prompt injection jailbreak testing data-exfiltration testing privilege/access-control testing adversarial testing
- Security Frameworks/Tools: OWASP LLM guidance Garak PyRIT Promptfoo or equivalent AI red-teaming tools
- Observability: tracing and evaluation of prompts model responses tool calls agent decisions latency token usage and failures
- Degree in Computer Science or related field or equivalent professional experience.
- 6-8 years of experience in QA SDET software engineering or test automation.
- Strong understanding of QA methodologies test architecture and software development lifecycle.
- Strong coding/scripting skills in Python; Bash JavaScript or C experience is beneficial.
- Experience building automation frameworks and integrating automated tests into CI/CD pipelines.
- Strong Linux command-line debugging and troubleshooting skills.
- Experience with API functional integration regression performance and security testing.
- Experience with Kubernetes containers and virtualization.
- Hands-on exposure to LLMs Generative AI applications RAG AI agents or Agentic AI Platforms.
- Understanding of LLM-specific failure modes including hallucination non-determinism prompt injection and unsafe agent behavior.
- Experience with Selenium Playwright PyTest or similar testing frameworks.
- Knowledge of HPC/AI infrastructure networking GPUs or distributed systems is a plus.
Build a modern quality engineering practice where traditional software testing and AI evaluation come together enabling us to ship infrastructure software and AI agents that are reliable measurable secure and safe to operate in production.
Location
Bangalore
Travel
-
Required Experience:
IC
About Company
Penguin Solutions designs, builds, deploys, and manages large, complex Al and high-performance computing (HPC) infrastructures at scale.